Phonon-2 ships: a 164MB open English ASR that beats Whisper

Fermion's open-weight model transcribes an hour of audio in 20 seconds on a MacBook, runs fully offline, and drops into any Whisper or OpenAI STT pipeline.

Nowline OCT 7 2:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 5.21% WER, 164 MB, tops Whisper turbo

    Phonon-2 averages 5.21% word error across seven English test sets — LibriSpeech, AMI, Earnings-22, GigaSpeech and more — versus 6.58% for Whisper large-v3-turbo, in a file that downloads at just 164 MB. Fermion bills it as the most accurate open English ASR under 900 MB.

  • An hour of audio in 20 seconds, offline

    On an M5 MacBook Air it runs 174x realtime via MLX (one hour of speech to text in ~20 seconds), 40x on CPU alone, and up to 6,680x batched on a single H100 — all local, nothing leaving the machine. Transcription you used to rent from an API now costs nothing per minute.

  • Drop-in OpenAI-compatible server

    pip install fermion-research, then `fermion serve` starts an OpenAI-compatible /v1/audio/transcriptions endpoint — point your existing Whisper or OpenAI STT code at it and it just works. `fermion listen` does live mic transcription, and CPU and CUDA Docker images ship for servers.

  • Check the license before you ship

    Phonon-2's weights are CC-BY-4.0 (inherited from NVIDIA's Parakeet TDT 0.6B v3), so commercial use needs attribution. The CLI and the older Phonon-1 and Phonon-1-Micro models are Apache-2.0 — reach for those if attribution is a dealbreaker.

  • Build this weekend: private voice features

    A 164 MB offline model makes on-device dictation, meeting notes, podcast search or voice commands viable with no per-minute bill and no audio uploaded anywhere — the privacy-sensitive projects a cloud STT API tends to price or policy out of reach.