DeepSeek V4 Flash 0731 goes open weights under MIT, agentic leap

Same 284B/13B MoE and $0.14 pricing as the preview, but Terminal Bench leaps to 82.7 and it lands top-3 among open-weights models. Day-0 GGUFs are up.

Nowline AUG 2 6:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • MIT weights, self-host today

    DeepSeek pushed the full V4 Flash 0731 weights to Hugging Face under the MIT license — unrestricted commercial use, self-hostable via vLLM or SGLang. It's the official release that supersedes July's preview, so you can run the exact model behind the API on your own boxes.

  • Agent scores nearly doubled

    It's a re-post-train, not a new architecture, aimed squarely at agents: Terminal Bench 2.1 jumps from 61.8 to 82.7 and DeepSWE from 7.3 to 54.4. The Artificial Analysis Intelligence Index climbed 40 to 50, putting it among the top 3 open-weights models.

  • Still 13B active, still $0.14

    The refresh keeps the 284B-total / 13B-active MoE and 1M-token context, and API pricing holds at $0.14/$0.28 per million tokens with cache hits at $0.003 (~98% off). A speculative-decoding module ships attached for faster local inference.

  • Build this weekend: an agent you fully own

    Unsloth's day-0 GGUFs are already live, so a quantized 13B-active model with 1M context fits on a single workstation. Native Responses API support plus low/high/max reasoning make it a drop-in backend for a coding or terminal agent — no vendor in the loop.

  • The catch: verbose, and pricier than it looks

    Flash is chatty — it spent ~230M output tokens running the Intelligence Index versus a ~92M median, so real task bills land above the headline rate. DeepSeek also has peak-hour pricing (all rates double for 7 hours a day) planned, though not yet switched on.