DeepSeek ends the price war: V4-Pro API rates jump up to 1,100%

V4-Pro-0813 is stable at 1M context and SWE-bench 80.6 - but the cheap-inference era ends: cache tokens cost ~12x more, off-peak the only discount left.

Nowline AUG 17 11:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Cache-hit tokens took the biggest hit

    Cached input rocketed from $0.0036 to $0.044 per 1M tokens - roughly +1,100% - while peak uncached input tripled to $1.32 and output leapt to $3.96 from $0.87. If your agents replay a cached system prompt every turn, that line item just went up about 12x.

  • Off-peak scheduling is the new discount

    Live since 16:00 UTC on Aug 16, pricing now splits peak vs off-peak, with off-peak at half price: $0.66 input / $1.98 output. Even that discounted output still runs ~2.3x the old flat $0.87 - batch non-urgent jobs into the off-peak window or watch costs climb.

  • This ends China's inference price war

    DeepSeek's floor-scraping rates set the pace that dragged rivals down all year; Caixin frames this hike as a deliberate break from that war. Plan for Chinese-model inference to trend up, not down - the era of near-free frontier tokens is closing.

  • Your Responses API code now runs on DeepSeek

    V4-Pro and V4-Flash now natively speak the OpenAI Responses API format, so an existing Responses-based agent can switch base URL and keys and run unchanged. Both models also gained low/high/max thinking-effort levels to trade latency for depth per call.

  • The model that's charging the premium

    The GA 0813 checkpoint posts SWE-bench Verified 80.6 (near Opus 4.7's 80.8), LiveCodeBench 93.5, and a 1M-token context holding >95% retrieval to ~900K. It's a genuine frontier coder - the new prices just mean weighing it against Qwen and Gemini Flash more carefully.