Sakana ships Fugu-Cyber, an agentic security model you must apply for

Reportedly matches GPT-5.5-Cyber on the hardest vuln benchmarks — but the scores are self-reported, access is manually vetted, and it's blocked in the EU.

Nowline JUL 26 12:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • It orchestrates specialists, it isn't one big model

    Fugu-Cyber is an endpoint in Sakana's Fugu family that reads a security query, spins up specialist sub-agents, and runs a Thinker–Worker–Verifier loop so candidate vulnerabilities get checked before you see them. Sakana's pitch is that for security, orchestration beats raw model scale.

  • The benchmark numbers — and the asterisk

    It reports 86.9% on CyberGym (1,507 real bugs across 188 OSS-Fuzz projects) and 72.1% on Microsoft's CTI-REALM, edging GPT-5.5-Cyber's 85.6% and Claude Mythos Preview's 83.1%. Every figure is self-reported and un-replicated, and the methodology hasn't been published — treat the ranking as a claim until someone runs it independently.

  • You have to apply, and you can't self-host

    There are no open weights: access is API-only, every request is manually reviewed against a defensive-use policy, and it's unavailable in the EU/EEA. Budget for approval delays — there's no same-day evaluation if you want it in a workflow.

  • Pricing, and the 272K-token cliff

    It runs $6 per million input tokens and $36 output ($0.60 cached), across $20/$100/$200 Token Plans — but rates double once a request crosses 272K tokens, roughly a 20% premium over Fugu-Ultra. Cheap for targeted triage prompts, pricey for whole-repo sweeps.

  • What you'd actually build with it

    The real unlocks are detection engineering — turning threat reports into rules via MITRE ATT&CK mapping and KQL iteration — and vulnerability triage with proof-of-concept generation. Keep a human in the loop: Sakana itself says the model alone isn't enough for production security.