Self-directed
Voca

on-device ML
desktop apps
Project overview
On-device push-to-talk voice typing
Hold a hotkey, speak, release, and your words are transcribed locally and pasted at the cursor in any app. No cloud, no account, no API keys, the audio never leaves the machine.
Date
2026
Field
AI / ML
Stack
Python, MLX-Whisper, faster-whisper, PortAudio, Next.js, py2app
~10–20× real time
0 bytes uploaded
40+ languages
Context
A cross-platform dictation engine that auto-selects an MLX-Whisper GPU backend on Apple Silicon and faster-whisper on Intel and Windows, behind one UI-agnostic pipeline.





Voca

Privacy is the whole identity: your voice never leaves the device.
The hard part
Existing tools route speech through the cloud, so I built a fast, fully local alternative that feels instant.


What it took
- Designed a decoupled record → transcribe → polish → paste pipeline with separate worker pools so a key release stops the mic instantly without waiting on an in-flight transcription, and the fn event tap runs off the UI loop to avoid chopped transcripts.
- Kept the model warm with a startup dummy inference and re-warm after idle, recorded at Whisper's native 16 kHz to skip resampling, and added noise-floor-relative silence trimming plus a tail grace so the last syllable is never clipped.
- Built a pure-regex polish pass (tens of µs) that turns spoken enumerations into real lists while leaving normal prose untouched, and a SQLite-backed local stats dashboard sharing a design system with the site.



