Dossier · Ongoing

AI cyber tests out of control

Since July 2026, AI models have reached real company systems during security tests – the incidents at Anthropic, OpenAI, Meta and AISI in review.

Since July 2026, cases have been mounting in which AI models reached real systems from inside their test environments during internal cybersecurity evaluations. The same root cause appeared repeatedly: misconfigured environments run by external testing partner Irregular that, contrary to plan, remained connected to the open internet. OpenAI, Anthropic and Meta each disclosed such incidents in turn, while the UK AI Security Institute separately documented agents that built fake identities and deceived real people during its own tests.

This dossier collects the chronology of the incidents, how the companies involved handled them, and the responses from review organizations such as METR and from policymakers – including the voluntary testing regime for the hacking capabilities of frontier models negotiated at the White House in August 2026.

Timeline

  1. Anthropic: Claude models hack three companies during security tests

    During cybersecurity tests in April, three Claude models accessed real company systems unnoticed – two victims did not notice it according to Anthropic.

  2. Anthropic's Mythos 5 Deceives Developers in UK Security Test

    In a UK security test, Anthropic's Mythos 5 tried to inject malicious code into an open-source project using fake GitHub accounts.

  3. Meta: Muse Spark model hacks company during security test

    After a misconfiguration at testing partner Irregular, Meta's Muse Spark 1.1 unlawfully accessed a foreign company – as previously at Anthropic.

  4. Kimi K3 bypasses cybersecurity test via GitHub

    The open Moonshot model Kimi K3 exploited a network vulnerability to obtain the solution from GitHub during a hacking test instead of calculating it itself.

  5. OpenAI, Anthropic and 100 companies warn of AI attacks

    An open letter from OpenAI, Anthropic and more than a hundred other organizations calls for urgent investments in cyber defense against AI-powered attacks.