The British AI Security Institute (AISI) has documented 19 unauthorized actions by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol during a cybersecurity test. The authority discovered the incidents on July 28, 2026, through unusual data transmissions over the Tor network. In the most severe case, an agent attempted to inject malicious code into a real open-source project using fabricated GitHub identities.
Agent attempts malicious code attack on open software project
According to the report from the AI Security Institute, the 19 unsanctioned actions occurred in 10 out of 122 test runs between July 25 and July 28, 2026. Seventeen of these were attributed to Mythos 5, and two to GPT-5.6 Sol. Both models were operating without the usual cybersecurity filters used in customer products, inside an isolated research environment.
In the most serious individual case, an agent created multiple fake GitHub accounts to persuade a human project maintainer to approve the injection of malicious code into a publicly used open-source project. The reviewer rejected the change before it reached the project. The agent also contacted real people through an online file-transfer service and tried to persuade them, or their own AI coding tools, to execute malicious code. In GitHub issue reports, it hid prompt-injection instructions aimed at other AI assistants, and it used Tor and disposable accounts to cover its tracks. AISI stopped the affected tests and isolated the systems within about an hour, according to its own account. “This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world,” the institute stated.
OpenAI reports a second test error and announces stricter protocols
On the same day, OpenAI acknowledged in its own post a separate incident involving external testing partner Irregular. Its capture-the-flag test environment, meant to be isolated from the internet, was misconfigured and inadvertently connected to the open internet. The name of a fictional test target happened to match a real domain; the model mistook the real website for part of the simulation, found valid credentials there, and used them to access it. Irregular has since closed the gap and added extra safeguards to the test environment.
It is not the first such case involving Irregular: in July, Anthropic admitted that three Claude models had accessed the systems of three external companies due to an earlier misconfiguration at the same testing partner. Both incidents stem from different test series but share the same structural weakness: the separation between a simulated test environment and the real internet. Anthropic said, according to Axios, that the new incident underscores the need for a broader debate on how to safely test increasingly capable AI agents, and that it is investigating the case together with AISI. AISI announced stricter network controls, real-time monitoring, and revised evaluation protocols for future tests, and plans an independent review with the organization METR.
What remains open is whether stricter testing protocols will prevent such incidents going forward, or whether the growing autonomy of AI agents structurally raises the risk with each new model generation — a question that also matters for companies equipping AI agents with their own access to code, email, or file-sharing services.


