Update: GPT-6 Astra runs your computer for 40-min autonomous tasks

Yesterday's API model now drives apps end-to-end: 72.6% on OSWorld, two zero-days found in testing, and access widening to Plus, Pro and the API.

Nowline SEP 4 11:00 PM banner

Top AI stories from the last hour

Top AI stories from the last hour

Copy markdown

  • It drives real software now, not just chat

    Astra completed multi-step jobs — converting graphics to 3D, building a game, editing a contract, booking a service — by operating the apps itself. On the OSWorld computer-use benchmark it hit 72.6% (up from Sol's 65.7%) while cutting average task time from ~75 to about 40 minutes.

  • The benchmark sweep, in numbers

    FrontierMath Tier 4 97.6%, ARC-AGI-3 98.6% (with the Responses API harness), DeepSWE v1.1 74.1% on agentic coding, BenchCAD Vision2Code 95.9%, and Mind2Web tasks completed 1.9x faster than the Sol-based setup. The practical read: fewer babysitting loops on long, tool-heavy runs.

  • Access is widening past the enterprise gate

    Day one was limited to OpenAI's Enterprise Daybreak partners; the rollout to Plus, Pro, Business and API customers starts "in the coming days," with rate-limit credits for the wait. If you were locked out yesterday, watch your dashboard this week.

  • An AI engineer for roughly $6 an hour

    At ~33 tokens/sec against the $50/M output rate, running Astra flat-out pencils out to about $6/hour for autonomous build-and-debug work. On coding agents it also burns roughly a third of the tokens rivals do, per Artificial Analysis — so the real bill can undercut the sticker rate.

  • It crossed OpenAI's 'Critical' cyber line

    Astra is the first model to trip the top cybersecurity tier in OpenAI's preparedness framework, and it surfaced two previously unknown software flaws during testing. Expect tighter guardrails on the API — and a good reason to keep your own agent runs sandboxed.

  • It holds its place across long sessions

    Astra takes notes across context windows, searches back through earlier messages, and can ask clarifying questions without pausing its other parallel work — the coherence that multi-hour, fleet-of-subagents agent jobs have been missing.