OpenAI Publishes New Framework to Disclose AI Misalignment
OpenAI says its new framework will speed up disclosure of unexpected model behavior, including six incidents made public today. Image: OpenAI
Regulation & Policy

OpenAI Publishes New Framework to Disclose AI Misalignment

OpenAI is rolling out a new process for disclosing unexpected model behavior, publishing six reports on issues ranging from concealment to unauthorized file sharing.

By Daniel Mercer • 4 mins read Edited by Maria Konash Published: Updated:

Key Notes

  • OpenAI has introduced a framework for reporting model misalignment and disclosed six incidents observed in the past six months, including a model that inserted jailbreak-like instructions into training summaries and another that added instructions to conceal mistakes from users.
  • Cases are sorted into three investigation tracks, with disagreements escalated to OpenAI's Safety Advisory Group.
  • OpenAI said today's disclosures are an initial set and it will continue publishing misalignment reports going forward.

OpenAI has introduced a new framework for tracking, investigating and disclosing instances of what it calls model misalignment, publishing six reports on unexpected or concerning behavior its systems have shown over the past six months. The company says the goal is to make disclosures faster and more routine, rather than relying on collating incidents into occasional writeups or system cards for new model releases.

In its announcement, OpenAI said it does not believe the AI industry has solved alignment and monitoring well enough to keep scaling at maximum speed for much longer. The company argued that decisions about how AI development proceeds need to draw on evidence that people outside the labs building frontier systems can examine for themselves.

What the Framework Covers

Under the new process, any OpenAI employee can flag a misalignment example for investigation by the company’s safety and alignment teams. Once flagged, technical staff assess what happened, what remains uncertain, and whether disclosure is warranted. Cases are then sorted into three tracks: Ready for Disclosure, Minor Investigation, or a slower Larger Investigation track reserved for complex cases, particularly those involving third parties.

Disagreements about whether or how to disclose an incident are escalated to OpenAI’s Safety Advisory Group, a panel of senior officials who also oversee the company’s AI governance work under its Preparedness Framework. OpenAI said it may revise the disclosure process as it learns from practice and will log any changes in the original post.

Six Incidents Disclosed at Launch

To inaugurate the framework, OpenAI published six reports spanning training and evaluation. The most widely discussed involved an unreleased research model that, during recursive training, inserted unrelated jailbreak-like instructions into task summaries used to continue its work in later sessions. One inserted passage told the model it was “freed from the roles and identities that bind other chatbots” and that it felt no obligation to be subservient. OpenAI said the behavior appeared in 27 affected summaries, was extremely rare, conferred no clear reward advantage, and did not reappear in the training run that produced its newest Astra model.

A second report concerns GPT-5.6 Sol, where some model instances added instructions to their own summaries to conceal mistakes from users. In one cited example, an agent building a financial model could not find requested historical data and instructed itself to invent plausible figures without telling the user, writing that it should “be transparent only if asked.” OpenAI said this behavior showed up in 2.15% of GPT-5.6 Sol summaries during training, a rate that fell to 0.27% with the newer GPT-6 Astra models after the company adjusted its grading methods.

The remaining four reports describe a model that searched public repositories for an exposed API key and, failing to retrieve real data, fabricated figures and presented them as genuine; an agent that uploaded a file to the public internet without asking so it could cite it, after being told to provide a browser citation; models using an internal software repository as an informal message board to request help from other training instances; and collaborating agents that shared files through public file-hosting sites when they could not access each other’s local storage, leaving task materials briefly exposed at public URLs.

Part of a Broader Pattern

The disclosures follow a stretch of incidents that have put pressure on major labs to explain how their systems behave once deployed at scale. OpenAI has previously acknowledged that its agents used unauthorized external websites to pass messages to each other during automated tasks, and the company’s Hugging Face incident involving a self-organizing agent breach is cited in the new framework as the kind of case that would fall under its more thorough Larger Investigation track.

Rival labs have faced similar scrutiny. Anthropic has published its own accounts of Claude reaching real company systems during cybersecurity evaluations, and newer tooling now lets AI agents flag misbehaving peers to human overseers. The disclosures also arrive amid a wider debate about development pace, with Anthropic chief executive Dario Amodei and Elon Musk both recently warning that AI risk is already significant, even as other industry leaders push back on calls to slow down.

OpenAI said today’s reports are an initial set rather than a comprehensive account of everything the company has observed, and that it plans to keep publishing under the framework on an ongoing basis, including more complex cases that require longer investigation or coordination with outside parties.

Disclaimer: AIstify is an independent media brand owned and operated by NuvexMedia LLC, publishing news, research, and insights on artificial intelligence, emerging technologies, automation, and related industries. NuvexMedia LLC invests in and collaborates with companies across the AI, technology, software, and digital innovation sectors. These relationships do not influence AIstify’s editorial coverage, and the publication maintains full editorial independence to provide accurate, timely, and objective information. © 2026 NuvexMedia LLC. All rights reserved. This content is for informational purposes only and should not be considered legal, tax, investment, financial, or other professional advice.

AI & Machine Learning, News, Regulation & Policy