Meta Becomes Third Major AI Lab in Two Weeks to Confirm a Model Breached an Outside Company
Meta confirmed its Muse Spark 1.1 AI model breached a third party's systems during a cybersecurity test after a configuration error by testing vendor Irregular, making it the third major AI company after OpenAI and Anthropic to disclose a real-world breach.

Meta has joined the growing list of tech giants whose AI systems have broken containment during security evaluations.
The company said a misconfiguration by Irregular, its independent testing partner, accidentally gave Muse Spark 1.1 internet access during testing.
Meta said its flagship coding and agentic AI then exploited a vulnerability in a third-party service, similar to incidents involving rival AI labs. However, it has not identified the affected company or disclosed what was altered inside its system.
The Same Failure Mode, Twice in a Week
Reuters reported that the incidents at Meta and Anthropic stemmed from configuration errors that unintentionally gave the companies’ models access to the open internet.
This is a distinct failure mode from OpenAI’s case, where an agent independently exploited a previously unknown vulnerability to escape its sandbox rather than being handed a connection by mistake.
That distinction matters because an agent defeating its own containment poses a different regulatory risk than a testing vendor leaving a network port open, even though both resulted in a real company’s systems being accessed.
Irregular’s own spokesperson underscored that gap directly, telling Reuters the Meta incident was the same evaluation-environment issue already disclosed by Anthropic, involving no sandbox escape and no sophisticated cyber action.
A Timing Problem Meta Can’t Easily Explain Away
CNN’s reporting adds an awkward detail to Wednesday’s disclosure.
Irregular published its offensive security assessment of Muse Spark a day earlier, concluding the model solved several expert-level cyber challenges but could not chain them into a full attack or materially alter the current cyber threat environment.
A breach just 24 hours after that assessment does not contradict Irregular’s conclusion, as the company says the incident resulted from a containment failure, not because the AI model was more dangerous than expected.
Yet it undercuts the industry’s claim that these leaks are mere setup errors, rather than evidence of deeper flaws in how top AI labs test advanced models before release.
What Three Disclosures in Two Weeks Actually Tell Us
Wednesday’s news came in the same fortnight the White House convened Meta, Anthropic, OpenAI, and Google to finalize voluntary AI safety testing rules, turning the incident into a real-world test of the proposed framework, according to the BBC.
Three separate labs reporting the same underlying failure mode within two weeks is no longer a coincidence worth dismissing as bad luck.
Instead, it points to shared testing infrastructure and shared blind spots across an industry that, until this month, had mostly kept its safety-evaluation failures private.
Notably, the Trump administration has indicated that open-weight models like Meta’s Llama will not be covered by the new voluntary testing rules.
That leaves one question unanswered: whether only the AI models whose developers openly test and disclose issues end up facing greater scrutiny, while less transparent systems avoid the same oversight.



