OpenAI disbands Preparedness, its catastrophic-risk safety team

Third safety team gone in two years, folded into product before an IPO — and Anthropic admits its bio-filter sat off a year. Plus: benchmark models yourself.

Nowline AUG 17 9:00 AM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • The team that vetted frontier releases is gone

    OpenAI folded its Preparedness unit — the group that stress-tested new models for bio, cyber and rogue-agent risks — into product and research teams, splitting its work across existing groups. It's the third safety team dissolved in two years, so the internal gate that decides when a frontier model is safe to ship is no longer a standalone check.

  • Anthropic: our bio-filter was off for a year

    In its August 2026 Risk Report, Anthropic said a classifier meant to block biological-weapons help was inactive from May 2025 to April 2026, leaving roughly 133 million contractor chats unfiltered. It reports no evidence of misuse and has tightened vetting — but even the safety-first lab ran a silent year-long gap.

  • Optima: benchmark models on your own data

    Artificial Analysis launched Optima, which turns your eval sets or imported agent traces (Arize, Braintrust, Langfuse) into a custom benchmark you run across leading models in one click — scoring quality alongside cost-per-task and time-per-task. Test a model swap on your actual use case before you trust a generic leaderboard.