Cactus Whistle: a 16.9MB speech-to-text model that runs on CPU

On-device voice with no GPU and no API bill — plus a leaked Gemini mode that wants your whole Mac, and a DeepSeek desktop harness.

Nowline Oct 4 4:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 16.9MB, on the CPU, no dependencies

    Cactus's Whistle hits the first token in 11ms and decodes 1,319 tokens/sec — roughly 8.6x smaller and 6x faster than Whisper base, with lower word-error rates on LibriSpeech, SPGISpeech and Earnings-22. `pip install cactus-needle` then `needle.transcribe("clip.wav")`; weights are on Hugging Face and the engine deploys to 17 targets from microcontrollers to the browser. On-device dictation with zero cloud round-trip is now a weekend project.

  • Report: Gemini eyes full Mac access

    TestingCatalog surfaced an unreleased "Additional sandbox options" in Gemini Desktop that would reportedly let Gemini read, create, modify or delete files anywhere on a Mac and drive apps like Mail and Safari without asking first. Google hasn't confirmed it, and buying, accounts and legal terms would still need approval — but it's a reminder to scope exactly what you let an agent touch.

  • DeepSeek Harness v0.2 ships desktop installers

    The DeepSeek Harness agent added desktop builds for macOS, Windows and Linux with plugin management and workspaces — a local-first way to run DeepSeek models for coding and everyday work without living in a terminal.