DeepSeek V4-Pro promo pricing ends today — off-peak billing arrives

Output leaps 4.5x and input 3x. But you can halve the bill by shifting jobs to cheaper hours, self-host the MIT weights, or route around it entirely.

Nowline AUG 16 11:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Output up 4.5x, input up 3x

    DeepSeek's promotional V4-Pro rates expire today. At peak hours, cache-miss input goes from $0.435 to $1.32 per 1M tokens and output from $0.87 to $3.96 — roughly 3x and 4.5x. Even cached input jumps, from $0.0036 to $0.044 per 1M.

  • Off-peak billing halves it

    New this cycle: peak/off-peak pricing. During off-peak hours (10:00–01:00 UTC), rates drop to half of peak — output is $1.98 per 1M instead of $3.96. Shift batch jobs, evals, and async agents into that window and cut your DeepSeek spend by 50%.

  • The escape hatch is MIT-licensed

    V4-Pro's weights (V4-Pro-0813, a 1.6T-param MoE with 49B active and a 1M-token context) sit on Hugging Face under MIT. If the new API math breaks your budget, self-host or rent GPUs and skip per-token pricing entirely.

  • Routers haven't repriced yet

    The hike lands on DeepSeek's first-party API. OpenRouter still lists V4-Pro near the old promo (~$0.40 in / $0.79 out per 1M). Until third-party providers follow, they may be the cheaper front door to the exact same model.

  • Elsewhere: Grok 4.6 lands in GitHub Copilot

    xAI's Grok 4.6, tuned for long-running agent work, is now selectable in GitHub Copilot's model picker (Aug 14) — another frontier option alongside Claude and GPT right in your editor.