Cheap models by default, premium on call: how US teams cut AI bills
DeepSeek works at $0.87/M vs Fable's $50, and Chinese models now take 58% of US OpenRouter traffic — plus the routing recipes teams actually ship.

Copy markdown
The 57x price gap driving it
A million output tokens runs $0.87 on DeepSeek V4-Pro and $4.40 on Z.ai's GLM-5.2, against $15 for Kimi K3 and roughly $50 for Anthropic's Fable. That spread is what teams now architect around instead of eating.
The routing recipe, spelled out
The Journal reports Telnyx swapped premium Anthropic across ~1,400 agents for Z.ai, keeping Fable only for planning and OpenAI's Sol for review — cheap open weights do the build, premium models plan and check. Harvey ships a button that calls Fable 5 only when a task genuinely looks hard.
The receipt that sells it
Cursor reportedly clocked a from-scratch browser build at about $10,000 on GPT-5.5 alone versus $1,339 on a mixed Composer + Opus 4.8 stack — roughly a 7x cut for comparable output. The premium model becomes the specialist, not the workhorse.
How mainstream this already is
Chinese models took 58% of US firms' tokens on OpenRouter as of July 20, peaking near 63% earlier in the month, up from under 10% in early 2025; the US share has slid from ~80% to ~42%. IDC pegs 47% of US firms with 1,000+ staff using a Chinese model for at least one use case.
What to do this week
Meter your traffic and route bulk and default calls to a cheap open-weight tier — DeepSeek V4, GLM-5.2, Qwen — while reserving premium models for planning and final review. Before moving production, read each model's license and data-residency terms: the savings are real, but so are the compliance footguns.