Update: open-weight Kimi K3 tops Arena's coding leaderboard

The 2.8T model now beats Opus 4.8 and GPT-5.6 Sol on frontend code, ranks #1 open-weight on Agent Arena, and runs at $3/$15 per Mtok or self-hosted.

Nowline JUL 27 7:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • First open-weight model to lead a code arena

    Kimi K3 climbed to #1 on Arena's frontend-coding board within hours of release — reportedly 1,679 points — edging Claude Opus 4.8 and GPT-5.6 Sol. It's the first self-hostable model to top a public coding leaderboard.

  • The weights are actually downloadable now

    Moonshot's promised July 27 drop landed: the full 2.8-trillion-parameter weights are live on Hugging Face. Full precision still wants 64+ Blackwell-class GPUs, so unless you have a rack, the hosted API is the realistic path.

  • Frontier agentic coding at $3/$15

    Call K3 through Moonshot's API at $3/M input ($0.30 cached) and $15/M output with a 1M-token context — well under Opus 5's $5/$25 — or hit the same INT4 endpoint via OpenRouter and Cloudflare Workers AI.

  • #1 on task success, but weak on steering

    On Agent Arena it lands #4 overall, level with Opus 4.8 and GPT-5.6 Sol, and #1 on confirmed task-success across 8K+ live sessions. The catch: it trails the field on steerability and ships with an undisclosed hallucination rate.

  • What you can build with it

    A whole-repo coding agent that matches closed frontier models on frontend work — feed it the full 1M-token context for cross-file refactors, and if you self-host, your per-token cost drops to electricity.