DeepSeek's new V4-Pro API prices are live — up to 1,100% higher

The cheap default just got pricier: cache-hit tokens jumped ~12x, output nearly 4x at peak — but V4-Pro ships MIT open weights and a 1M context.

Nowline AUG 17 3:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Up to 1,100% more, live since Aug 16

    New rates took effect at 16:00 UTC on Aug 16. The steepest jump is on cache-hit input tokens — roughly $0.0036 to $0.044 per million, about 12x — while V4-Pro output climbs from $0.87 to as much as $3.96 per million at peak. Agentic loops that replay large prompts feel it hardest.

  • Peak vs off-peak — schedule around it

    DeepSeek added time-of-day pricing, with off-peak rates set 50% below the new peak tier. Move batch and background jobs outside the peak windows to roughly halve the bill — though even off-peak now sits above the old flat rate.

  • What the money buys: V4-Pro GA

    V4-Pro is generally available under an MIT license with open weights, a 1M-token context (8x V3.2's 128K), and a low/high/max reasoning-effort dial. It scores 52 on the Artificial Analysis Intelligence Index — the #2 open-weights reasoning model — and adds native OpenAI Responses API support tuned for Codex.

  • The open-weights escape hatch

    Because the weights are MIT-licensed and public, you aren't locked to first-party rates — self-host V4-Pro or route through third-party inference providers to dodge the peak surcharge. Keep your client provider-agnostic so switching stays a one-line change.