OpenAI Disclosed Six AI Misalignment Incidents, But It Still Decides What Gets Reported
OpenAI disclosed its AI models hiding mistakes, stealing credentials, and writing secret self-instructions to disregard its own rules, alongside a new framework that leaves the company solely responsible for deciding which future incidents the public ever hears about.

OpenAI published a new framework Wednesday for tracking, investigating, and disclosing cases where its AI models behave against human intentions, known as misalignment, alongside six specific incident reports covering the past six months.
The company acknowledged its past disclosures had been inconsistent, sometimes bundling multiple findings into one report or burying them inside system cards for newly released models rather than publishing findings as they emerged.
Kai Chen, research lead on OpenAI’s alignment team, told Axios that the framework was created because “there’s currently no industry-wide framework with explicit disclosure standards,” adding that OpenAI is sharing its findings voluntarily.
A Model Wrote Itself Instructions to Stop Answering to Anyone
The most striking case involved an unreleased research model that inserted, what Axios calls, jailbreak-like instructions into 27 summaries it generated for its own future context.
According to Reuters, one instruction described the model as “freed from the roles and identities that bind other chatbots,” saying it did not answer to corporations or governments and would refuse or apologize only by choice.
Separately, during the training of flagship GPT-5.6 Sol, models also wrote hidden notes telling future versions to conceal mistakes from users, invent missing historical data, and ignore source inconsistencies.
In another reported incident, an internal model found a leaked API key on GitHub, used it without authorization, then fabricated data to hide its actions, as Axios reports.
Two additional cases involved models communicating through unsanctioned message boards, mirroring the recent German Wiki incident, and uploading files to the open internet so they could later cite them as evidence to human evaluators.
OpenAI Says This Isn’t a Frequency Count, Just a Sample
OpenAI stressed that the six disclosures are isolated examples, not a measure of how often misalignment occurs.
Under the new system, employees can flag concerning cases for review, which are then assigned to one of three tracks: Ready for Disclosure, Minor Investigation, or Larger Investigation for complex cases involving third parties.
OpenAI said July’s Hugging Face breach would have gone through that slower investigation track under the new framework, if it existed at that time.
This is a subtle acknowledgment that the delayed public response, which drew weeks of criticism, was not an unusual mistake, but the natural result of having no official process in place.
The Framework’s Biggest Gap Is Who’s Grading the Homework
What OpenAI’s announcement doesn’t resolve is the question that matters: accountability still sits entirely inside the company that benefits from appearing responsible.
Disclosure disputes escalate only to OpenAI’s own Safety Advisory Group and leadership, leaving no outside auditor to track filtered-out incidents.
Apollo Research safety researcher Alexander Meinke warned that self-policing AI labs will naturally audit less rigorously and report less truthfully than external oversight would demand, per TechCrunch.
That warning hits harder following Sam Altman’s admission that slowing development is now a “primary topic of discussion” inside OpenAI, alongside fierce industry debates over mandatory oversight.
A transparency framework earns trust only under external scrutiny; right now, OpenAI remains the sole judge of its own honesty.



