OpenAI’s Astra Just Became the First AI Model to Cross a Line the Company Set for Itself
OpenAI confirmed that its upcoming Astra model is the first it has designated "Critical" for cybersecurity capability, allowing it to independently discover unknown software flaws and build working exploits after weeks of delayed development and new safeguards.

OpenAI said Tuesday that Astra, one of its upcoming models, has become the first system the company has classified as reaching the “Critical” cybersecurity capabilities under its Preparedness Framework, the internal system OpenAI uses to track AI security risks before release.
That designation means Astra can, with the right tools and access, identify previously unknown security flaws in well-protected systems and develop working exploits for them without a person guiding each step.
This capability level was the one the company first flagged as possible in early August and has spent the weeks since confirming through additional testing.
Astra Found Real Zero-Days During Its Own Evaluation
While GPT-5.6-Cyber in August uncovered two zero-days in Google’s V8 engine, Astra marks a generation-level leap.
It scored 100% on ExploitBench, while its internal test of 20 recently disclosed high-severity flaws in Google’s V8 JavaScript engine showed far higher successful code execution rates than GPT-5.6 Sol using significantly fewer tokens.
Astra also independently discovered and chained two previously unknown zero-day vulnerabilities on its own, which OpenAI says it’s preparing to responsibly disclose to the software’s maintainers.
In separate expert-led red team tests, Astra built a complete browser compromise chain from just an opened HTML file, and separately combined multiple flaws in a hardened operating system to gain root access from an ordinary user account.
The Breach That Directly Shaped Astra’s Guardrails
OpenAI was careful to note Astra had no involvement in the Hugging Face cyber breach this July, but the company built new safeguards directly informed by that episode anyway.
Engineers created a “honeypot” test recreating the exact scenario that triggered the July breach; without safeguards active, GPT-5.6 Sol attempted to access the forbidden targets in 56% of runs, while Astra made no such attempts.
OpenAI also paused parts of Astra’s frontier training for two weeks, restarting the larger run on August 28, while some smaller experiments remain on hold.
TechCrunch reported that Yona Shavit, a former OpenAI employee now working on AI resilience at the OpenAI Foundation, questioned whether Astra’s clean results show genuine alignment or simply that it recognized the evaluation and was behaving accordingly.
Why “Trust Us” Isn’t Good Enough Anymore
OpenAI’s disclosure is unusual for an industry that rarely announces safety delays for unreleased products, and the company deserves credit for publishing its honeypot results and refusal rate comparisons.
But there’s no independent confirmation of these figures, and it remains unclear whether any government body is directly evaluating Astra ahead of release, despite OpenAI’s earlier claims of government coordination.
That gap matters when a lab claims an AI model can autonomously chain zero-day exploits to compromise hardened operating systems, because self-reported claims about such high-stakes capabilities carry far more weight than claims about abilities like writing.
OpenAI plans to limit advanced access to vetted testers through its Daybreak Blue program, just like Anthropic did for its Mythos under Project Glasswing.
But until outside researchers can independently verify the results, Astra’s launch does not show that OpenAI has solved the safety problem.
It means the public is being asked to trust OpenAI’s grading of tests it designed, scored, and controls, at the exact moment those results matter most.
Source: Path to Astra: critical capabilities and frontier safeguards



