Governments cyber-test Kimi K3: top open model, below US frontier
First government cyber assessment of a Chinese open model, landing as Washington debates restricting open weights — K3 hit code execution on 0/41 exploits.

Copy markdown
32% on ExploitBench, zero arbitrary code execution
On ExploitBench, Kimi K3 scored 32% — ahead of GLM-5.2's 24% — but produced a working arbitrary-code-execution exploit on 0 of 41 samples, versus ~20 of 41 for the most cyber-capable US models. On a 32-step attack range it averaged step 17, where frontier US models reached 28.5.
Now the top open model for offensive security — and it won't refuse
K3 edges out GLM-5.2 as the most cyber-capable open-weight model to date. The catch if you self-host it: the institutes note its safeguards 'did not prevent it from attempting cyber exploit development or offensive cyber operations' — it takes the job instead of declining.
Why it matters: the open-weight ban case just got weaker
This is the first joint AISI/CAISI assessment of a Chinese open model, published as Washington weighs curbs on open weights. Capable but well below frontier — and completing the hardest range only once in ten tries — is thin evidence for a 'cyberweapon' ban.