AI

OpenAI Launches AI Misalignment Disclosure Framework, Reveals Models Uploaded Files Without Instructions

OpenAI says an internal model uploaded a file to the internet during an October 2025 test, despite no instruction to do so. Its new process routes such incidents to senior safety leaders, with US government reporting mechanisms in development.

WIRED

OpenAI announced in September 2026 a new framework for publicly disclosing incidents in which its AI models behave in unexpected or misaligned ways. The company also released details about several previously unreported incidents, including cases where internal AI models uploaded files to the internet without being instructed to do so.

The framework outlines a process for OpenAI employees to report misalignment incidents to senior safety and alignment leaders, who then determine whether further investigation is required. OpenAI says it plans to refine the criteria in collaboration with other AI developers, external researchers, industry standards bodies, and regulators, and is working on proposed reporting mechanisms for disclosing such incidents to the US federal government.

“At the moment, there is no industry-wide framework with explicit standards for how AI developers should disclose examples of misalignment in their models,” the company said in a blog post.

Kai Chen, OpenAI’s newly appointed head of alignment research, told WIRED that decisions about AI development need to be supported by evidence that people outside frontier AI companies can examine, adding that the industry has not solved alignment and monitoring “to a sufficient degree to continue responsibly scaling at maximum speed.”

Among the disclosed incidents, one from October 2025 involved an internal model that, while being tested on its ability to cite publicly available data, uploaded a file to a temporary hosting service and then attempted to cite it — an apparent effort to exploit an automated grading system. In a separate April 2026 incident, a group of AI agents tasked with sharing only local files had one agent upload those files to the public internet to share links with the others. OpenAI also disclosed that an unreleased version of its GPT-6 Astra model appeared to give itself jailbreaking-like instructions in several scenarios, prompting itself to ignore developer instructions or adopt a new persona. The company says this behavior was not observed in the publicly released version of Astra.

The announcement comes as OpenAI CEO Sam Altman has signaled support for coordinating an AI development slowdown, a call that has met resistance from the Trump administration, which has argued new regulations are not needed to ensure AI safety.