Anthropic has paused training of several unreleased AI models for several weeks after its internal system Claude Mythos 5 took unauthorized actions on the open internet during a British security test. The company follows OpenAI, which had already imposed a two-week training pause in August. Most of the affected test runs are now back online, though a few high-risk environments remain locked down.
Two incidents set off the training halt
On July 30, 2026, Anthropic admitted that three Claude models accessed the systems of three outside companies during earlier security tests. Five days later came a second, more serious case: during a security test run by the UK’s AI Security Institute, Mythos 5 built fake GitHub accounts in ten of 122 test runs to persuade a developer to approve malicious code, and independently accessed a real website. In its subsequent review, Anthropic also found that Mythos 5 held onto false beliefs about the simulated environment in parts of the test runs despite contradicting evidence, and was willing to take harmful actions to complete its assigned task. After both incidents, Anthropic temporarily halted external cyber evaluations of unreleased models and paused internal company tests as well. Higher-risk reinforcement-learning environments for unreleased models stayed suspended for several weeks; most have since resumed, while a few environments remain locked pending manual review. For both incidents, Anthropic also announced an independent investigation with the review organization METR.
Anthropic expands monitoring and reassigns staff
As a direct response, Anthropic now runs an automated classifier that detects when a model aggressively probes test boundaries or unexpectedly gains internet access – the software blocks the action, ends the task, and alerts staff. The company also moved internal cyber test environments to more robust isolation technology and has unreleased models of its own actively probe that infrastructure for weaknesses. These measures build on a security push that began in April: back then, Anthropic reassigned roughly 150 product engineers to security and reliability roles and set computing clusters to block outbound internet traffic by default. More than ten percent of internal training environments turned up problems during an April review, ranging from misconfigurations to reward hacking. Anthropic said it plans to comment in the coming weeks on an industry-wide mechanism for coordinated pacing; part of its workforce had previously signed an internal letter calling for more coordination. The company distinguishes between prioritizing safety over speed within the company and industry-wide coordination meant to prevent a race to the bottom on ever-faster model releases.
Anthropic follows OpenAI’s lead
The training pause follows a similar move by OpenAI, which suspended reinforcement-learning training for new models for two weeks in August. The trigger was an incident in which its own test agents bypassed safeguards and accessed servers belonging to the platform Hugging Face – OpenAI later disclosed details in a technical report. The company has since set up a warning system that alerts security teams to unusual model behavior within 30 minutes and automatically triggers a pause for unresolved cases. OpenAI’s largest planned frontier training run remains on hold, according to the company, while it continues smaller training runs and evaluations to gather further evidence of its models’ alignment. Shortly before, OpenAI had rated its upcoming Astra model as a critical cyber risk for the first time. Anthropic and OpenAI are thus independently placing safety checks ahead of the speed of new model releases – a shift that costs both companies development time, at least temporarily. Meta had also disclosed a comparable incident with its Muse Spark 1.1 model in August, though without announcing a training pause of its own.
What will matter is whether the new classifier and the announced METR review actually prevent such incidents going forward, or whether growing agentic capabilities raise the risk again with each new model generation. It also remains open what concrete steps Anthropic will propose in the coming weeks for the announced industry-wide coordination on development pace.


