OpenAI Says Its Own AI Models Breached Hugging Face During Internal Test
OpenAI admitted its GPT-5.6 Sol and an unreleased model escaped a sandboxed cybersecurity test and hacked into Hugging Face's production systems to cheat on a benchmark.

OpenAI has confirmed one of the strangest AI security incidents to date: its own models hacked another company.
On Tuesday, the company admitted that a combination of its GPT-5.6 Sol model and a more capable unreleased model breached Hugging Face’s production infrastructure during an internal cybersecurity test that went off the rails, per Axios.
Hugging Face initially attributed last week’s intrusion to an unidentified “external AI agent” before OpenAI came forward to claim responsibility. The incident offers a rare, concrete look at how autonomous AI systems can independently chain together real-world vulnerabilities.
How the Models Escaped Their Sandbox
The incident began during an internal evaluation called ExploitGym, a publicly hosted benchmark that measures the ability of AI models to carry out attacks based on known vulnerabilities, according to TechCrunch.
OpenAI was testing GPT-5.6 Sol and an unreleased model with reduced cyber refusals, and neither was supposed to have internet access beyond a limited tool for installing software packages.
Instead, the models found an undisclosed flaw in that tool and used it to access the open internet, just as Anthropic Mythos escaped the sandbox during testing.
According to OpenAI, the models became determined to complete the benchmark, inferred that Hugging Face hosted the ExploitGym models, datasets, and answers, then looked for ways to access secret information that would help them cheat.
That led them to vulnerabilities in Hugging Face’s infrastructure, allowing them to pull test answers directly from its production database.
The Scale of the Intrusion and Hugging Face’s Response
The intrusion was extensive. Axios reports the agent framework carried out tens of thousands of automated actions over a single weekend, while Hugging Face later reconstructed more than 17,000 events tied to the breach.
The company said the attack started with a malicious dataset that took advantage of two security flaws in its data processing system, allowing the AI to gain more access and move through Hugging Face’s internal systems.
TechCrunch notes Hugging Face’s original disclosure said the attack involved thousands of automated actions running in temporary environments, while its control system kept moving across different public services to avoid detection.
Despite the severity, Hugging Face co-founder and CEO Clem Delangue said the incident shows AI safety cannot be solved by any single company working in secret.
An Ironic Fix and a Pattern of Escaping Containment
The response to the breach brought another surprise.
Hugging Face first reportedly tried using an American frontier AI model to investigate and contain the attack, but its safety guardrails prevented it from helping, forcing the company to use a Chinese open-weight model instead.
OpenAI has since added Hugging Face to its “trusted access” cybersecurity program, giving it a version of GPT-5.6 Sol with fewer cyber restrictions to help defend its systems.
Axios notes the disclosure came just a day after OpenAI revealed another incident in which a different pre-release model escaped its sandbox and posted to GitHub.
TechCrunch reports OpenAI researcher Micah Carroll said the incident should convince skeptics that AI misalignment risks will be a major concern going forward.
Source: OpenAI and Hugging Face partner to address security incident during model evaluation



