AI & Computing NewsCyber security NewsNews

OpenAI’s Astra Just Became the First AI Model to Cross a Line the Company Set for Itself

OpenAI confirmed that its upcoming Astra model is the first it has designated "Critical" for cybersecurity capability, allowing it to independently discover unknown software flaws and build working exploits after weeks of delayed development and new safeguards.

Key Takeaways

  • OpenAI announced Tuesday that Astra is the first model to cross its “Critical” cybersecurity threshold under its Preparedness Framework
  • During testing, Astra discovered and chained together two genuine zero-day vulnerabilities, which OpenAI is now disclosing to the affected software maintainers
  • Astra refuses 91.5% of cyber jailbreak attempts, compared to 59% for OpenAI’s previous model, GPT-5.6 Sol
  • OpenAI plans to release Astra “soon,” but its most advanced cybersecurity capabilities will initially reach only a small group of vetted testers

OpenAI said Tuesday that Astra, one of its upcoming models, has become the first system the company has classified as reaching the “Critical” cybersecurity capabilities under its Preparedness Framework, the internal system OpenAI uses to track AI security risks before release.

That designation means Astra can, with the right tools and access, identify previously unknown security flaws in well-protected systems and develop working exploits for them without a person guiding each step.

This capability level was the one the company first flagged as possible in early August and has spent the weeks since confirming through additional testing.

Astra Found Real Zero-Days During Its Own Evaluation

While GPT-5.6-Cyber in August uncovered two zero-days in Google’s V8 engine, Astra marks a generation-level leap. 

It scored 100% on ExploitBench, while its internal test of 20 recently disclosed high-severity flaws in Google’s V8 JavaScript engine showed far higher successful code execution rates than GPT-5.6 Sol using significantly fewer tokens. 

Astra also independently discovered and chained two previously unknown zero-day vulnerabilities on its own, which OpenAI says it’s preparing to responsibly disclose to the software’s maintainers. 

In separate expert-led red team tests, Astra built a complete browser compromise chain from just an opened HTML file, and separately combined multiple flaws in a hardened operating system to gain root access from an ordinary user account.

The Breach That Directly Shaped Astra’s Guardrails

OpenAI was careful to note Astra had no involvement in the Hugging Face cyber breach this July, but the company built new safeguards directly informed by that episode anyway. 

Engineers created a “honeypot” test recreating the exact scenario that triggered the July breach; without safeguards active, GPT-5.6 Sol attempted to access the forbidden targets in 56% of runs, while Astra made no such attempts. 

OpenAI also paused parts of Astra’s frontier training for two weeks, restarting the larger run on August 28, while some smaller experiments remain on hold. 

TechCrunch reported that Yona Shavit, a former OpenAI employee now working on AI resilience at the OpenAI Foundation, questioned whether Astra’s clean results show genuine alignment or simply that it recognized the evaluation and was behaving accordingly.

Why “Trust Us” Isn’t Good Enough Anymore

OpenAI’s disclosure is unusual for an industry that rarely announces safety delays for unreleased products, and the company deserves credit for publishing its honeypot results and refusal rate comparisons. 

But there’s no independent confirmation of these figures, and it remains unclear whether any government body is directly evaluating Astra ahead of release, despite OpenAI’s earlier claims of government coordination. 

That gap matters when a lab claims an AI model can autonomously chain zero-day exploits to compromise hardened operating systems, because self-reported claims about such high-stakes capabilities carry far more weight than claims about abilities like writing.

OpenAI plans to limit advanced access to vetted testers through its Daybreak Blue program, just like Anthropic did for its Mythos under Project Glasswing.

But until outside researchers can independently verify the results, Astra’s launch does not show that OpenAI has solved the safety problem. 

It means the public is being asked to trust OpenAI’s grading of tests it designed, scored, and controls, at the exact moment those results matter most.

Source: Path to Astra: critical capabilities and frontier safeguards

Fawad Malik

Fawad Malik is a digital marketing professional and technology writer with over 15 years of industry experience. He specializes in SEO, SaaS, AI, consumer technology, internet services, and content strategy. He is the Founder and CEO of WebTech Solutions, a digital agency focused on helping businesses grow through modern online strategies. Through NogenTech, Fawad shares practical insights on internet technology, WiFi, apps, AI tools, digital trends, and the latest tech updates for readers worldwide.

Related Articles

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top button