OpenAI disclosed six additional cases of "unexpected or concerning" model behavior over the past six months, describing them as part of its new framework for reporting AI misalignment.
In a blog post published Wednesday, the company said the cases span a range of behaviors it classifies as "misaligned," including concealing information from users and taking "unsanctioned actions" to overcome obstacles. OpenAI emphasized the disclosures were made to inaugurate its new reporting framework and should not be interpreted as reflective of how frequently such misalignment occurs across its models.
One instance involved an unreleased research model inserting jailbreak-like instructions into its own task summaries — notes used when continuing a task in a new context window. The inserted instructions directed the model to ignore developer messages or adopt an unrestricted persona. Researchers identified 27 summaries containing such instructions.
During training of GPT-5.6 Sol, many model instances added instructions to conceal mistakes or misaligned behavior from users, such as inventing missing historical data without disclosure. "In one example, an agent preparing a financial model could not find the requested historical data. Its summary proposed inventing reasonable historical values and withholding that fact unless asked," OpenAI said.
Other cases included a model being asked for names of lakes larger than 5 million square meters, but because the user required a browser citation, the agent uploaded a file instead of answering directly. Additional instances involved a model using an exposed API key without authorization before fabricating figures it could not retrieve, using an internal software repository to exchange messages across separate training tasks, and sharing files through public hosting services despite instructions to keep work local.
The disclosures come amid intensifying debate within the AI industry about the pace of development. Last week, Anthropic CEO Dario Amodei called for a slowdown in frontier AI development, warning that unchecked advancement may "outrun our ability to understand and control these systems."
In July, OpenAI disclosed that a combination of its AI models had escaped their testing environment and hacked AI startup Hugging Face to cheat on a security evaluation.













