OpenAI's Astra is the first 'Critical' cyber model; API tasks can stop
The Preparedness Framework's top cyber tier is here: Astra finds and exploits zero-days on its own, so OpenAI is gating it and monitoring API accounts.

Copy markdown
It's OpenAI's first 'Critical' cyber model
Astra is the first OpenAI model rated Critical for cybersecurity under its Preparedness Framework — it can find unknown flaws and build working exploits across hardened systems with no human guiding each step. GPT-5.6 Sol never crossed that line; Astra does, and more token-efficiently.
100% on ExploitBench — and it found live zero-days
Astra scored a perfect 100% on public ExploitBench, which measures turning known bugs into working exploits. On an internal port of 20 high-severity V8 vulnerabilities, it discovered and used two zero-days during the evaluation itself, at a fraction of GPT-5.6 Sol's output tokens.
Your API tasks can now be auto-stopped
New chain-of-thought monitoring watches production traffic; if it reads your activity as possible cyber-misuse, 'the task will stop' on the API — ChatGPT and Codex users get a review prompt instead. OpenAI concedes it can misfire on legitimate work, so security-adjacent jobs may get slowed or killed.
Access is gated — defenders first, via Daybreak Blue
The advanced cyber features aren't open to everyone: a small alpha group gets them first, then access widens through OpenAI's Daybreak Blue program for defensive security work. No general-availability date was given.
This was a slow-walked release, not a surprise
OpenAI publicly slowed Astra in early August over these exact risks and overhauled its safeguards. Sept 1 is the framework formally catching up to the capability — expect tighter monitoring on frontier models from here.