Inkling-Small weights drop: 80% on SWE-bench, runs on two H200s

The open-weights coder beats its 975B sibling on SWE-bench, runs on two H200s in 4-bit, and reads text, image and audio — fine-tune it on Tinker now.

Nowline AUG 3 11:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Open weights, and it out-codes its 975B sibling

    276B total, 12B active. Inkling-Small posts 80.2% on SWE-bench Verified and 31.6% on HLE — past the full 975B Inkling on both — and ships as open weights (reported Apache 2.0) you can run commercially.

  • Self-host it on two H200s

    The NVFP4 checkpoint fits in ~180GB: W4A4 on a single B300, or W4A16 across two H200s. It runs on vLLM, SGLang, Unsloth and Hugging Face today — a near-frontier coder with no per-token bill.

  • Text, image and audio in, 1M context

    Native multimodal input — images as 40x40 patches, 16kHz audio clips — with a 1M-token window and 90.1% on VoiceBench. One open model for your doc, screenshot and voice pipelines.

  • Fine-tune it tonight on Tinker

    Both Inkling models are on Tinker at a limited-time discount for text, image and audio chat plus fine-tuning; the weights also reach Together, Fireworks, Modal, Databricks and Baseten. Tuned for cheap coding, grading and synthetic-data jobs.

  • The catch: it is shakier on facts

    SimpleQA Verified falls to 20.6% versus the big Inkling’s 43.9% — the small model hallucinates more closed-book. Bolt on retrieval before trusting it for knowledge work.

  • Elsewhere: Microsoft’s own full-duplex voice model surfaces

    MAI-Realtime reportedly appeared in a hidden Microsoft playground preview: a bidirectional voice model with two voices that can run web search and tools mid-conversation, hinting Copilot may drop its OpenAI-Realtime dependency.