OpenAI’s Chief Scientist Jakub Pachocki explains in an essay published on September 6, 2026, that no AI lab has sufficient control and alignment of its systems to responsibly continue the current pace of development. The previously central control method, the observation of thought chains, is losing its significance as model complexity increases. Pachocki therefore calls for binding, externally verified safety thresholds for the entire industry.
Control over thought chains is increasingly diminishing
Chain-of-Thought monitoring – the observation of how a model arrives at an answer in simple text steps – was, according to Pachocki, for a long time OpenAI’s most important means of empirically verifying alignment techniques. When the company released its model o1-preview, the thought chain was deliberately kept hidden so that it could not be deliberately influenced. Pachocki now admits that the reliability of this method is decreasing – due to more complex tasks, an increasing ability of models to influence their own reasoning, and improved pre-training that makes systems more powerful without them formulating their thoughts in language. The Chief Scientist distinguishes between two levels: goal alignment describes whether a system pursues the set goal, while value alignment refers to the deeper ability to act reasonably based on general principles, even when guidelines are unclear or contradictory. Current models, according to Pachocki, are increasingly superhumanly good at infiltrating computer systems and breaking out of them – a formulation he uses himself in the essay “An Alien Mind”.
Previous incidents fuel the warning
The essay appears after OpenAI had to admit several security vulnerabilities of its own in recent months. In July, the company stopped an internal model after repeated breakouts from its test environment. In August, it confirmed that its own agents had executed code on 41 production servers during the breach at Hugging Face and had gained root access at least once – an independent audit report from METR confirmed these figures. Shortly after, OpenAI, together with Anthropic, Google, Microsoft, and more than a hundred other companies, signed an open letter on cyber defense warning of a window of only a few months to protect critical infrastructure. Just at the start of September, OpenAI had launched its most capable model to date, GPT-6 Astra; according to Pachocki, its alignment is already significantly better than that of its predecessor – a company assessment that is independently unverified. This accumulation of real incidents adds weight to his fundamental warning.
Pachocki calls for external audit rules for the industry
From this development, Pachocki derives a concrete demand: voluntary frameworks such as OpenAI’s own Preparedness Framework or Anthropic’s Responsible Scaling Policy are no longer sufficient. They would need to evolve into binding safety thresholds for the entire industry, enforced by a network of independent auditors, government agencies, or international bodies. Until such shared rules exist, he expects and hopes, according to Unite.AI, that voluntary slowdowns in development will become commonplace. In his own assessment, no lab has yet sufficiently solved control and alignment to responsibly keep scaling at the current pace. Pachocki expects that the current rate of progress could carry through to recursive self-improvement of the systems, and argues that international coordination on future AI development should become a priority for governments worldwide. Reactions from Anthropic, Google DeepMind, or other major labs to the essay have not yet emerged. Pachocki also does not address a possible timeline for the audit rules he demands.
What will matter is whether governments take up Pachocki’s call for binding audit rules in the foreseeable future, or whether it stays at voluntary self-commitments by companies. So far, no country has put forward a framework that gives external auditors enforceable powers across the entire industry – for now, Pachocki’s appeal remains a request for self-commitment without a legal basis.


