Inkling-Small weights drop: 80% on SWE-bench, runs on two H200s
The open-weights coder beats its 975B sibling on SWE-bench, runs on two H200s in 4-bit, and reads text, image and audio — fine-tune it on Tinker now.

Copy markdown
Open weights, and it out-codes its 975B sibling
276B total, 12B active. Inkling-Small posts 80.2% on SWE-bench Verified and 31.6% on HLE — past the full 975B Inkling on both — and ships as open weights (reported Apache 2.0) you can run commercially.
Self-host it on two H200s
The NVFP4 checkpoint fits in ~180GB: W4A4 on a single B300, or W4A16 across two H200s. It runs on vLLM, SGLang, Unsloth and Hugging Face today — a near-frontier coder with no per-token bill.
Text, image and audio in, 1M context
Native multimodal input — images as 40x40 patches, 16kHz audio clips — with a 1M-token window and 90.1% on VoiceBench. One open model for your doc, screenshot and voice pipelines.
Fine-tune it tonight on Tinker
Both Inkling models are on Tinker at a limited-time discount for text, image and audio chat plus fine-tuning; the weights also reach Together, Fireworks, Modal, Databricks and Baseten. Tuned for cheap coding, grading and synthetic-data jobs.
The catch: it is shakier on facts
SimpleQA Verified falls to 20.6% versus the big Inkling’s 43.9% — the small model hallucinates more closed-book. Bolt on retrieval before trusting it for knowledge work.
Elsewhere: Microsoft’s own full-duplex voice model surfaces
MAI-Realtime reportedly appeared in a hidden Microsoft playground preview: a bidirectional voice model with two voices that can run web search and tools mid-conversation, hinting Copilot may drop its OpenAI-Realtime dependency.