OpenAI's Astra is the first 'Critical' cyber model; API tasks can stop

The Preparedness Framework's top cyber tier is here: Astra finds and exploits zero-days on its own, so OpenAI is gating it and monitoring API accounts.

Nowline SEP 2 2:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • It's OpenAI's first 'Critical' cyber model

    Astra is the first OpenAI model rated Critical for cybersecurity under its Preparedness Framework — it can find unknown flaws and build working exploits across hardened systems with no human guiding each step. GPT-5.6 Sol never crossed that line; Astra does, and more token-efficiently.

  • 100% on ExploitBench — and it found live zero-days

    Astra scored a perfect 100% on public ExploitBench, which measures turning known bugs into working exploits. On an internal port of 20 high-severity V8 vulnerabilities, it discovered and used two zero-days during the evaluation itself, at a fraction of GPT-5.6 Sol's output tokens.

  • Your API tasks can now be auto-stopped

    New chain-of-thought monitoring watches production traffic; if it reads your activity as possible cyber-misuse, 'the task will stop' on the API — ChatGPT and Codex users get a review prompt instead. OpenAI concedes it can misfire on legitimate work, so security-adjacent jobs may get slowed or killed.

  • Access is gated — defenders first, via Daybreak Blue

    The advanced cyber features aren't open to everyone: a small alpha group gets them first, then access widens through OpenAI's Daybreak Blue program for defensive security work. No general-availability date was given.

  • This was a slow-walked release, not a surprise

    OpenAI publicly slowed Astra in early August over these exact risks and overhauled its safeguards. Sept 1 is the framework formally catching up to the capability — expect tighter monitoring on frontier models from here.