Back to the stories

OpenAI designates GPT-6 Astra as first model with 'Critical' cybersecurity capability and gates advanced cyber features

Score 10.0,

Cybersecurity

Building on our coverage from the day before, OpenAI now says GPT-6 Astra is its first model to be rated "Critical" for cybersecurity, and it will keep the model's most powerful hacking capabilities tightly gated.

OpenAI introduced GPT-6 Astra on September 4, 2026 and is rolling it out in stages: early access for customers in its Daybreak cybersecurity program, then paid ChatGPT tiers, API users and cloud partners. Free-tier users and the smallest paid plans are excluded at launch.

Why this matters: Astra isn't just more capable in general tasks. In internal evaluations it scored 100% on ExploitBench and, in modified tests, identified and exploited two previously unknown vulnerabilities. That combination, frontier offensive capability plus real-world proof it can find novel bugs, changes the balance between powerful defensive tools and potential misuse.

How it works, in plain terms: think of Astra as a massively parallel red team. Instead of one analyst trying ideas in sequence, the model can try many attack paths at once and stitch them together. OpenAI says it trained Astra to refuse harmful requests and added monitoring that can interrupt suspicious behaviour, plus stronger sandboxing and controls on internet access.

What changes now: defenders gain a much more capable tool and OpenAI is funding access and training for critical operators with a $1 billion program, but advanced features are limited to vetted organisations and strict configurations. OpenAI also paused some frontier work, tightened internal safety processes, and says engineers are building automated shutdowns to stop runaway behaviour. Practical reality: using Astra's cyber mode requires enterprise-scale infrastructure and oversight, not a laptop experiment.

The open question is policy as much as tech: will these technical safeguards and gated access prevent misuse while letting legitimate defenders adopt Astra? That will determine whether this step is net progress for security or a new risk to manage.