JetBrains open-sources Mellum2.1, a 12B MoE coder you self-host

Apache-2.0 weights run quantized on one GPU. Plus DeepSeek 4.1 Flash’s cheap-coding surge, a 17MB open speech model, and Cloudflare’s open Clef.

Nowline OCT 10 11:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • A 12B coding agent you can run on one GPU

    JetBrains open-sourced Mellum2.1-12B-A2.5B-Thinking — 12B total, 2.5B active per token, 131K context, Apache 2.0. It posts 82.0 on LiveCodeBench v6 and runs as a Q4 GGUF in ~8.1GB, a local reasoning sub-agent that explores a repo, edits files, and runs tests. It still trails Qwen3.5-9B on the hardest SWE-bench Pro and Terminal-Bench tasks.

  • DeepSeek 4.1 Flash is the cheap coder builders flocked to

    A “why isn’t everyone freaking out” post topped Hacker News (860+ points): DeepSeek V4.1 Flash runs $0.016/M input, $1.20/M output, 1M context, native vision, with compressed KV cache cut to ~a quarter of the last Flash. Paired with flat-rate harnesses like OpenCode Go, builders report near-unlimited agent sessions that rarely top $1.

  • Whistle: speech-to-text in a 16.9MB file

    Cactus’s open ASR model is a single 16.9MB file that beats Whisper base on LibriSpeech and runs CPU-only — phones, wearables, even microcontrollers — at ~11ms to first token on an M4 Pro. `pip install cactus-needle`, then transcribe with word timestamps and keyword biasing. Weekend fuel for on-device voice.

  • Cloudflare’s open Clef decides in ~39ms

    Clef and Clef-flash are Apache-2.0 “decision models” (27B and 9B, Qwen backbones) that return typed answers with probabilities instead of prose — tool routing, triage, moderation, even image checks, up to 64 questions a call. Clef-flash’s median latency is ~39ms, and the API is drop-in compatible with TypeSafe’s Jev.