Kimi K3's 2.8T open weights drop at 00:00 UTC on July 27

The biggest open model yet leads Arena's code board — but serving it wants ~1.4TB of memory, its license is still unpublished, and it runs at one pricey tier.

Nowline JUL 26 8:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • Frontier scale, now downloadable

    Moonshot posts Kimi K3's weights to Hugging Face at 00:00 UTC July 27: a 2.8T-parameter sparse MoE (896 experts, 16 active per token, so ~50B effective compute) with a 1M-token context. It's the first frontier-tier open model at this scale — though Moonshot calls it a commitment, and open-weight dates have slipped before.

  • The catch: ~1.4TB of memory

    The MXFP4 4-bit weights are ~594GB on disk but need ~1.4TB of fast memory to actually serve — roughly eighteen 80GB accelerators, or a single 8×Blackwell/MI400 node that barely fits them. 16-bit runs ~5.6TB. This is a datacenter drop, not a laptop one.

  • It leads on code — at one expensive gear

    K3 tops Arena's Frontend/Code board (~1,679 Elo) and beats Claude Opus 4.8 max and GPT-5.5 high on knowledge work per Artificial Analysis, while trailing Claude Fable 5 and GPT-5.6 Sol. The snag: it exposes only a 'max' reasoning tier — Simon Willison's pelican SVG burned 13,241 tokens and 25¢.

  • License still unpublished

    Moonshot hasn't posted K3's license, so commercial use stays unconfirmed until it ships alongside the weights. K2 used a Modified MIT that required attribution once you passed 100M users or $20M revenue — budget for a similar string before you build on it.

  • Can't self-host? Rent it today

    If eighteen GPUs is out of reach, K3 already runs via the OpenAI-compatible API at platform.kimi.ai and free on the kimi.com web app. Self-hosting's main payoff is keeping inference traffic off Moonshot's servers — it doesn't cover training or the company's legal exposure.