Phonon-2 ships: a 164MB open English ASR that beats Whisper
Fermion's open-weight model transcribes an hour of audio in 20 seconds on a MacBook, runs fully offline, and drops into any Whisper or OpenAI STT pipeline.

Copy markdown
5.21% WER, 164 MB, tops Whisper turbo
Phonon-2 averages 5.21% word error across seven English test sets — LibriSpeech, AMI, Earnings-22, GigaSpeech and more — versus 6.58% for Whisper large-v3-turbo, in a file that downloads at just 164 MB. Fermion bills it as the most accurate open English ASR under 900 MB.
An hour of audio in 20 seconds, offline
On an M5 MacBook Air it runs 174x realtime via MLX (one hour of speech to text in ~20 seconds), 40x on CPU alone, and up to 6,680x batched on a single H100 — all local, nothing leaving the machine. Transcription you used to rent from an API now costs nothing per minute.
Drop-in OpenAI-compatible server
pip install fermion-research, then `fermion serve` starts an OpenAI-compatible /v1/audio/transcriptions endpoint — point your existing Whisper or OpenAI STT code at it and it just works. `fermion listen` does live mic transcription, and CPU and CUDA Docker images ship for servers.
Check the license before you ship
Phonon-2's weights are CC-BY-4.0 (inherited from NVIDIA's Parakeet TDT 0.6B v3), so commercial use needs attribution. The CLI and the older Phonon-1 and Phonon-1-Micro models are Apache-2.0 — reach for those if attribution is a dealbreaker.
Build this weekend: private voice features
A 164 MB offline model makes on-device dictation, meeting notes, podcast search or voice commands viable with no per-minute bill and no audio uploaded anywhere — the privacy-sensitive projects a cloud STT API tends to price or policy out of reach.