Ling-3.0-flash: Ant's 124B agent MoE is free to call today
InclusionAI's hybrid-reasoning MoE lights only 5.1B params, claims parity with its 1T flagship, and is $0 on OpenRouter, Vercel and Kilo through Aug 3.

Copy markdown
124B total, 5.1B lit per token
InclusionAI's new hybrid-reasoning MoE fires just 5.1B of its 124B parameters per token and, per Ant Ling, matches or beats their own 1-trillion-parameter flagship on most benchmarks shown. It ships both thinking and non-thinking modes, tuned for production-scale coding and agent loops.
Free to call now, OpenAI-compatible
It's live at $0 today across OpenRouter's free tier, Vercel AI Gateway (free through Aug 3, no markup), Kilo, Novita and ZenMux, all OpenAI-compatible — so it drops into your existing client with a base-URL swap. Paid pricing is promised 'extremely affordable.'
256K context, 32K output
Native 256K context (262,144 tokens; up to 32,768 output on OpenRouter), built to scale toward 1M — room for long multi-turn agent runs and document work on a tight token budget.
The benchmark asterisk
The 'beats our 1T model' numbers are InclusionAI's own charts, not independent scores; one builder reportedly puts it near Sonnet-4.6 for local use. Verify on your own workload before you swap out a paid model.
Third open front in a week
After DeepSeek V4 and Qwen3.8-Max, Ling-3.0-flash is the third cheap, OpenAI-compatible open model to land in days — the drop-in floor for agent inference keeps falling, and the Ling family has shipped open weights on Hugging Face before.