Update: DeepSeek V4 Pro's GA price triples at peak, live Aug 16
The $0.44 rate was the preview; GA adds peak/off-peak tiers where even off-peak costs more, plus a Codex-ready Responses API and Copilot model cuts.

Copy markdown
The $0.44 rate was the preview price
DeepSeek's headline bargain was the four-month preview rate; from Aug 16 at 16:00 UTC, V4 Pro GA moves to peak/off-peak tiers — off-peak runs about $0.66/M input and $1.98/M output, and peak roughly triples input to ~$1.32/M and output to ~$3.96/M. Off-peak is 50% of peak, but even that discounted tier sits above today's flat $0.435/$0.87, so budget for a real increase.
The cache discount collapsed — agents feel it most
The sharpest hit for agent loops: cache-hit input jumps from ~$0.0036/M to ~$0.022/M off-peak, shrinking the discount from roughly 1/120th of input to about 1/30th. Workflows that re-read the same files or system prompts every turn lose most of their caching savings — re-price your runs before Aug 16.
What you gain: a one-click Codex endpoint
It's not all cost: GA ships native OpenAI Responses API support “optimized for Codex with one-click setup,” plus thinking-effort tiers — low, high, and max — on both V4 Pro and V4 Flash. You can point Codex or any Responses-API client straight at DeepSeek without writing a translation shim.
Elsewhere: GitHub Copilot retires six models Sept 1
Prep your Copilot policies: on Sept 1, GitHub drops Claude Opus 4.5 and 4.6, Sonnet 4.5 and 4.6, Gemini 3.1 Pro, and Raptor Mini (annual individual plans keep Sonnet 4.6). Suggested swaps are Opus 4.7/4.8/5, Sonnet 5, and Gemini 3.6 Flash — and Gemini 3.7 Flash just landed in the picker.