IBM's Granite 4.2: open agentic reasoning models you run locally
3B–30B under Apache 2.0, a thinking toggle and 128K context, SWE-Bench 57 on the 30B — plus Granite Speech 5.0 and a 4-bit model that beats its original.

Copy markdown
Open weights, Apache 2.0, three sizes
Granite 4.2 lands in 3B, 8B, and 30B — all Apache 2.0, so you can download, fine-tune, and ship to production with no license strings. They're dense decoder-only models (not MoE), with a thinking / non-thinking toggle plus a low-effort mode that saves tokens on easy queries.
Agentic RL, with the numbers to back it
The 8B and 30B were trained with “agentic RL” in real sandboxes — using tools, writing and running code, searching the web — and speak OpenAI-format tool calls. The 30B posts SWE-Bench Verified 57.0, Terminal-Bench 2.1 29.2, and BFCL v4 61.4; the 8B hits 47.7 on SWE-Bench. A real local coding-agent brain, not a toy.
Runs on your box, 128K context
128K context (pretrained out to 512K), speculative decoding, and it runs on vLLM, SGLang, or Ollama — live today on Hugging Face, GitHub, OpenRouter, Replicate, and DeepInfra. Weekend build: a tool-using coding agent with zero per-token cost.
Granite Speech 5.0 Turbo, free transcription
IBM also shipped Granite Speech 5.0 Turbo CTC — 470M params, roughly 2x faster than the prior Open ASR Leaderboard leader, transcribing about three hours of audio per second. Apache-2.0 speech-to-text you can bolt onto anything, locally.
Elsewhere: a 4-bit model that beats its full-precision self
Multiverse Computing's “Quantization-Aware Healing” recovers — and slightly exceeds — full-precision accuracy after squeezing a model to 4-bit, with weights and the recipe on Hugging Face. Cheaper local inference without the usual quality tax.