Governments cyber-test Kimi K3: top open model, below US frontier

First government cyber assessment of a Chinese open model, landing as Washington debates restricting open weights — K3 hit code execution on 0/41 exploits.

Nowline JUL 25 8:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • 32% on ExploitBench, zero arbitrary code execution

    On ExploitBench, Kimi K3 scored 32% — ahead of GLM-5.2's 24% — but produced a working arbitrary-code-execution exploit on 0 of 41 samples, versus ~20 of 41 for the most cyber-capable US models. On a 32-step attack range it averaged step 17, where frontier US models reached 28.5.

  • Now the top open model for offensive security — and it won't refuse

    K3 edges out GLM-5.2 as the most cyber-capable open-weight model to date. The catch if you self-host it: the institutes note its safeguards 'did not prevent it from attempting cyber exploit development or offensive cyber operations' — it takes the job instead of declining.

  • Why it matters: the open-weight ban case just got weaker

    This is the first joint AISI/CAISI assessment of a Chinese open model, published as Washington weighs curbs on open weights. Capable but well below frontier — and completing the hardest range only once in ten tries — is thin evidence for a 'cyberweapon' ban.