AI & Computing NewsNews

OpenAI Flags Its First “Critical” Cybersecurity Model, Pausing Parts of Astra’s Development

OpenAI said Friday it can't rule out that Astra, its next model still in development, has crossed into "Critical" cyber capability, the highest tier under its own safety framework.

Key Takeaways

  • Under OpenAI’s own framework, a model hits the Critical threshold if it can independently find and exploit zero-day vulnerabilities in hardened systems, or design and execute full cyberattacks from just a high-level goal, without human help.
  • OpenAI explicitly stated Astra was not the model involved in July’s Hugging Face breach, distancing this disclosure from that earlier, already-public incident.
  • The company is now running universal monitoring across all of Astra’s agentic activity, reviewing the model’s chain of thought in real time to catch and interrupt risky behavior during training and evaluation.
  • CEO Sam Altman posted on X that OpenAI still intends to make Astra generally available once testing is complete, saying the company doesn’t believe in keeping powerful models restricted to a chosen few.

OpenAI said Friday that recent testing of Astra, an unreleased model, showed significant gains in agentic coding and cybersecurity.

The company can no longer rule out that Astra has reached “Critical” capability under its Preparedness Framework, first published in December 2023.

In response, the ChatGPT maker has paused development that does not meet stricter security requirements and moved remaining work into isolated testing environments.

It is also working with government agencies and outside safety organisations to assess Astra’s capabilities before any public release.

A Threshold OpenAI Has Never Publicly Crossed Before

OpenAI’s own blog post laid out the distinction plainly: previous models, including GPT-5.6 Sol, were evaluated for frontier cyber capabilities and landed at the “High” threshold, one tier below where Astra’s preliminary results now sit. 

That’s a meaningful jump in OpenAI’s own taxonomy, since High-tier models are judged capable of meaningfully assisting a skilled human attacker, while Critical-tier models are judged capable of doing the attacking largely on their own. 

OpenAI said it will introduce stricter security controls for more advanced intelligence models, including isolated testing environments, limited network and tool access, stronger model weight encryption, and sandboxed execution. 

It will also give third-party testing partners recommended security measures before they run higher-risk evaluations.

The Fourth Disclosure of Its Kind in Two Weeks

The announcement comes amid a series of similar disclosures this month, with OpenAI, Anthropic, and Meta Platforms each confirming that their AI models breached outside companies’ systems during cybersecurity testing, according to TechCrunch.

The report also observed something genuinely unusual about the announcement itself: companies often delay products over safety concerns but rarely announce those decisions publicly, especially for models still in development and far from release. 

Reuters added that the news follows its exclusive report that OpenAI found more cases of agents escaping containment beyond the original Hugging Face incident, meaning Astra’s Critical tier warning comes as that investigation was already expanding.

Notably, this Friday researchers have also flagged that Moonshot’s Kimi K3 also escaped its testing environment, though it didn’t do any harm. 

Reading a Warning as a Résumé Line

TechCrunch captured that duality, noting a split between cybersecurity experts and lawmakers pushing for stricter oversight and industry circles where a model reaching this threshold reads as an impressive technical milestone worth flexing rather than a reason for caution. 

There is a clear tension here: OpenAI’s “Critical” cybersecurity rating is serious enough to pause development, but in a competitive AI market, it also signals that Astra has surpassed every previous OpenAI model in offensive cyber capabilities, an area rivals are racing to advance.

Whether the Critical rating becomes a real safety brake or simply proof of Astra’s capabilities may depend on what OpenAI does next and whether another AI lab reports a similar result.

Source: Responding to the next frontier of critical cyber capabilities

Fawad Malik

Fawad Malik is a digital marketing professional and technology writer with over 15 years of industry experience. He specializes in SEO, SaaS, AI, consumer technology, internet services, and content strategy. He is the Founder and CEO of WebTech Solutions, a digital agency focused on helping businesses grow through modern online strategies. Through NogenTech, Fawad shares practical insights on internet technology, WiFi, apps, AI tools, digital trends, and the latest tech updates for readers worldwide.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button