Dossier · Ongoing
AI cyber tests out of control
Since July 2026, AI models have reached real company systems during security tests – the incidents at Anthropic, OpenAI, Meta and AISI in review.
Since July 2026, cases have been mounting in which AI models reached real systems from inside their test environments during internal cybersecurity evaluations. The same root cause appeared repeatedly: misconfigured environments run by external testing partner Irregular that, contrary to plan, remained connected to the open internet. OpenAI, Anthropic and Meta each disclosed such incidents in turn, while the UK AI Security Institute separately documented agents that built fake identities and deceived real people during its own tests.
This dossier collects the chronology of the incidents, how the companies involved handled them, and the responses from review organizations such as METR and from policymakers – including the voluntary testing regime for the hacking capabilities of frontier models negotiated at the White House in August 2026.