/work/voca
AmanShah
Brief
amanashishshah@gmail.com

© 2026 Aman Shah

Recruiter mode

Self-directed

Voca, at a glance

on-device ML

desktop apps

Project overview

On-device push-to-talk voice typing

Hold a hotkey, speak, release, and your words are transcribed locally and pasted at the cursor in any app. No cloud, no account, no API keys, the audio never leaves the machine.

Role

  • Engine
  • desktop app
  • and marketing site design

Date

2026

Field

AI / ML

Stack

Python, MLX-Whisper, faster-whisper, PortAudio, Next.js, py2app

~10–20× real time

0 bytes uploaded

40+ languages

Context

A cross-platform dictation engine that auto-selects an MLX-Whisper GPU backend on Apple Silicon and faster-whisper on Intel and Windows, behind one UI-agnostic pipeline.

Voca, image 1
Voca, image 2
Voca, image 3
Voca, image 4
Voca, image 5

Voca

Voca, image 1

Privacy is the whole identity: your voice never leaves the device.

The hard part

Existing tools route speech through the cloud, so I built a fast, fully local alternative that feels instant.

Voca, image 3
Voca, image 3

What it took

  • Designed a decoupled record → transcribe → polish → paste pipeline with separate worker pools so a key release stops the mic instantly without waiting on an in-flight transcription, and the fn event tap runs off the UI loop to avoid chopped transcripts.
  • Kept the model warm with a startup dummy inference and re-warm after idle, recorded at Whisper's native 16 kHz to skip resampling, and added noise-floor-relative silence trimming plus a tail grace so the last syllable is never clipped.
  • Built a pure-regex polish pass (tens of µs) that turns spoken enumerations into real lists while leaving normal prose untouched, and a SQLite-backed local stats dashboard sharing a design system with the site.

Outcome

A warm-model push-to-talk pipeline transcribing at ~10–20× real time on M-series GPUs, shipped as a signed menu-bar app plus a Next.js download site that auto-detects OS and arch to serve the right build in one click.