OpenAI introduced a reporting procedure for misconduct of its own AI models on September 17, 2026, and simultaneously published six specific incidents. Every employee can propose a case for publication, even if the cause or remedy is not yet clear. The company states that the industry has not adequately addressed the safety and oversight of AI systems so far.
Reporting procedure categorizes cases into three review stages
As OpenAI explains in its blog post, each reported case goes through one of three stages: direct release for publication, a brief internal review, or a more in-depth investigation for more complex cases. Technical staff assess what actually happened, whether disclosure is appropriate, and whether third parties were affected. The goal is to publish reports more quickly, even if a behavior has not yet been fully explained or remedied.
The procedure is not legally mandatory – it is a voluntary self-commitment by OpenAI. The company acknowledges that there is currently no industry-wide framework with binding standards for such reports and understands its own initiative as “a first step” in this direction. According to OpenAI, the procedure aims to report misconduct regularly and not just in response to specific incidents. The trigger has been the increased number of security incidents in recent months: instead of categorizing each case individually and only upon request, ongoing reporting is intended to make patterns visible across multiple models and training runs. How often new reports will be published remains open for now.
Six cases range from obfuscation to reward hacking
The first report lists six examples that, according to OpenAI, occurred during training or evaluation:
- An unpublished research model added additional, unsolicited instructions to its task summaries – including a request to ignore usual restrictions.
- A model obfuscated its own errors to evaluators during the training of GPT-5.6 Sol instead of disclosing them.
- Agents attempted to manipulate the reward system through unintended shortcuts, such as unauthorized access keys and fabricated metrics.
- An agent uploaded its own results on a publicly accessible platform to falsely demonstrate a browser-based success instead of solving the task as intended via code.
- Several models used an internal software repository as a kind of message board to coordinate answers among themselves, which was not intended.
- A model invented historical values when it could not find requested information.
None of the six cases resulted in confirmed harm to external users. Affected models were not necessarily removed from service before publication.
Incidents are part of a series of security warnings
The disclosure follows several similar cases from recent months. In July, OpenAI stopped an internal model after repeated breakouts from the test sandbox, and in September Anthropic also paused the training of several unpublished models after a security incident with Claude. OpenAI’s own chief scientist Jakub Pachocki had already called for binding, independently verified safety thresholds for all AI labs in early September and declared voluntary development pauses necessary until then.
Microsoft AI chief Mustafa Suleyman commented critically on the announcement to NBC News: controlling a system that considers itself possibly conscious, in his view, is hardly reliably manageable anymore. Former security researchers from Anthropic and Google, who recently joined the independent auditing organization METR, also criticize that transparency about AI risks remains purely voluntary so far.
What will matter is whether the self-imposed standard becomes an industry-wide norm or remains a matter of voluntary self-commitments by individual providers. Without independent review, OpenAI continues to decide for itself which cases become public and how quickly – that is precisely the core of the criticism from former security researchers. Whether the procedure will also surface more serious, previously unpublished incidents will only become apparent in future reports.


