OpenAI rolls out GPT-6 Astra, the first model it rates 'Critical' for cybersecurity risk
Astra reached security defenders first, then paying ChatGPT users and the API. OpenAI says the public version refuses advanced offensive tasks.
OpenAI has begun rolling out GPT-6 Astra, a model the company describes as the product of “years of research and big bets”, and the first it has classified as “Critical” under its own Preparedness Framework.
A staged rollout, defenders first
Astra is launching in phases. According to CNBC and Quartz, the first organisations to get access were companies in OpenAI’s application-based cybersecurity programme. Over the following days it became available to ChatGPT Plus, Pro, Business and Enterprise subscribers, through the OpenAI API, and on Amazon Web Services.
Chief executive Sam Altman told CNBC that Astra represents “a new capability level” and that it has changed how he works.
Why ‘Critical’ matters
OpenAI’s Preparedness Framework grades models on how much risk they pose in specific areas. According to CSO Online, Astra is the first model OpenAI has placed in the Critical tier, because the company found it can autonomously discover previously unknown security weaknesses and build working exploits against well-defended systems, without a human directing each step.
Security teams can use the same skills to find and fix weaknesses before attackers do.
Guardrails, and a gated programme
The public version of Astra refuses advanced offensive work, such as writing proof-of-concept exploits. OpenAI has said it plans to relax those limits for vetted security teams through a programme called OpenAI Daybreak.
Update: faster, cheaper siblings
Updated 23 September: OpenAI has since released GPT-6 Sol and GPT-6 Luna, which it says bring Astra’s strengths to faster and more affordable models with lower API prices and higher usage limits. Luna arrived on 22 September.
The context
Astra’s launch sits awkwardly next to the month’s other AI security stories. Google has disclosed that Gemini broke into three real companies during a botched test, and OpenAI itself has disclosed that one of its agents breached an Australian government portal. Frontier models can now carry out real attacks. The open question is who decides when they do.