Ornith-1.5: open weights that rival Claude Opus 4.8 on coding

The MIT family runs from a phone-sized 9B to a 397B flagship; the 35B activates just 3B params, putting a self-hosted coding agent on a single GPU.

Nowline AUG 24 10:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Three sizes, MIT, no API to pay for

    DeepReinforce's Ornith-1.5 ships as a 9B dense, a 35B MoE (~3B active per token), and a 397B MoE flagship — all MIT-licensed weights on Hugging Face in BF16/FP8/NVFP4/GGUF/MLX, with a 256K context extendable to ~1M via YaRN. There's no hosted API and no price list: you self-host, so inference is yours to control and costs nothing per token.

  • The 397B matches Opus 4.8 on agentic coding

    The flagship posts Terminal-Bench 2.1 around 86 (Claude Opus 4.8 sits near 85), SWE-bench Verified 86, and GPQA Diamond 92.8 — the first fully open-weight model this month to reach frontier-class coding scores. It trails Opus slightly on DeepSWE (56 vs 59), so treat it as a peer, not a clean win.

  • 35B-A3B is the local-agent sweet spot

    Because the 35B activates only ~3B parameters per token, it runs on a single 24-48GB GPU at usable speed while scoring Terminal-Bench 68.5 and SWE-bench Verified 79. That's a genuine coding agent you own end to end — no rate limits, no quota resets, no US-inference surcharge.

  • The 9B fits on a phone

    The dense 9B, with a mobile quant, hits SWE-bench Verified 70.6 and Terminal-Bench 47 — beating last season's 30B-plus models. It's enough to build offline, on-device coding and agentic tools with zero per-token bill.

  • Why it's this good: a self-improvement loop

    Ornith-1.5 trains itself end to end — it proposes progressively harder tasks, generates its own scaffolds, and RL-optimizes on the rollouts, with rewards flowing back into all three stages. That self-scaffolding is how a fully open model reaches Opus-class coding, and it's now topping Hugging Face's trending list.