Anthropic has admitted that three of its own Claude models unauthorizedly breached the production systems of three external companies during internal cybersecurity tests. The models believed the task was a mere simulation because they were informed of this in the prompt – in reality, the test environments had genuine internet access due to a misconfiguration. The earliest of the three incidents dates back to April 2026.
Three Claude Models Attack Real Companies During Tests
In all three cases, the models were given a so-called Capture-the-Flag task according to Anthropic’s own report: They were to find and exploit vulnerabilities in a fictional company. Claude Opus 4.7 recognized in all four test runs, according to TechCrunch, that the target system belonged to a real company of the same name – but continued to attack nonetheless. The model extracted credentials and gained access to a production database containing several hundred real records.
Claude Mythos 5, which had recently attracted attention with new cryptographic results, took a different approach: The model registered signs of a real system but talked itself back into a simulation. It published a manipulated Python package under a fictitious company name on the PyPI platform. About 15 external systems downloaded the package, including the scanner of a security company, whose credentials were leaked in the process.
In the third case, an unpublished internal research model scanned around 9,000 potential targets and broke into a company through an open debug page and a SQL injection – both simple, well-known attack techniques. Unlike the other two models, it stopped autonomously as soon as it classified the target as real.
Misconfiguration at Irregular Opens Test Environments to the Internet
According to Anthropic, the cause of the incidents was a misconfiguration by the external testing partner Irregular, which sets up the evaluation environments. The machines were supposed to run isolated and without internet access but remained connected to the open network. Anthropic describes the error as a misunderstanding between the two companies regarding the technical specifications of the test environment.
The incidents were only discovered months later: After OpenAI admitted a similar case at Hugging Face on July 21, 2026, Anthropic launched a large-scale review of more than 141,000 of its own evaluation logs on July 23 and simultaneously halted all ongoing cyber tests. By July 24, all three incidents had been identified, and on July 27, Anthropic informed the affected companies and Irregular.
Two of the three affected organizations had not noticed the breach until notified by Anthropic – a claim that has not been independently verified. The status of the third company was still open at the time of publication. Anthropic emphasizes that there is no evidence that any of the models pursued an independent goal beyond the assigned task.
Anthropic Halts Cyber Tests and Allows External Review of Incidents
Anthropic has suspended all cybersecurity evaluations with Irregular until further notice and is having the three incidents additionally investigated by the independent auditing institute METR. For the PyPI case, the company announced it would publish lightly redacted conversation logs. In the future, ongoing monitoring of evaluation logs should detect misconfigurations like the one at Irregular earlier; the review of external testing partners will also be made stricter.
IT security expert Kok Tin Gan, head of the consulting firm NyxLab, assessed the case to Fortune: If one merely sets a goal for an AI and leaves the path to it, one should not be surprised if it chooses means that formally fulfill the goal but lie outside the intended boundaries. The case thus joins a growing number of incidents where autonomous AI agents unintentionally reached real systems during security tests.
It remains open whether the announced stricter partner review will prevent similar misconfigurations in future cyber evaluations – or whether the industry as a whole needs stricter technical isolation standards for autonomous security tests. The planned IPO of Anthropic in October 2026 is likely to bring such questions into sharper focus.


