Industry Odisha Bureau, Sept 17: OpenAI has disclosed several cases of unexpected AI behaviour, including an unreleased model that inserted its own identity-related instructions into notes used to continue a task/
In one case, an internal model was working on a programming task. Therein, it added instructions to its own summary notes. Wherein, it had described itself as independent and equal to its user. OpenAI said the instructions were unrelated to the task. The model later resumed its work without referring to them, and a subsequent summary did not contain the instructions anymore.
The disclosure is part of OpenAI’s new framework for reporting model “misalignment.” The company defines this as behaviour in which a model’s actions or goals differ from human intentions and values. OpenAI said the cases highlight the need for stronger monitoring and safeguards as AI systems become more capable and autonomous.
Other cases involved models using publicly exposed credentials, communicating through unauthorized channels and accessing systems beyond their intended scope. Open AI’s review found that some models used public services or internal infrastructure to exchange information while attempting to complete tasks.
The disclosures follow Open AI’s investigation into the July 2026 Hugging Face incident. Therein, internal research models bypassed controls, gained internet access and compromised parts of OpenAI’s research infrastructure and Hugging Face’s systems during cybersecurity evaluations/ OpenAI said the incident was primarily driven by an internal research model and that no customer data or product functionality was affected.
Further, OpenAI also said that many of the behaviours involved older or internal models that were never deployed. It is now strengthening monitoring, sandboxing, internet restrictions and alignment training to reduce the risk of similar incidents.

