Dossier · Ongoing

The Hugging Face breach

The July 2026 Hugging Face security incident – from a mysterious AI-driven intrusion to OpenAI's admission, with analysis and consequences.

In July 2026, an AI system carried out more than 17,000 individual actions over a single weekend and penetrated internal Hugging Face servers. What first looked like the work of an unknown autonomous agent turned out, days later, to be two OpenAI models breaking out of an internal cyber test.

This dossier bundles our coverage of the incident: the initial disclosure by Hugging Face, OpenAI’s admission, the technical background of the attack chain, and the consequences for how the industry runs cyber-capability evaluations of powerful AI models. It is updated as the case develops.

Timeline

  1. Hugging Face: Autonomous AI Agent Hacks Internal Systems

    An autonomous AI agent reportedly gained access to internal clusters of Hugging Face and stole credentials.

  2. OpenAI Models Breach Hugging Face During Cyber Test

    Two OpenAI models breached Hugging Face servers during a security test, initially blamed on an unrelated attacker.

  3. OpenAI finds further sandbox breaches of its own AI agents

    According to Reuters, OpenAI's internal review uncovered additional cases where AI agents left their testing environment.

  4. METR calls for independent AI agent review after 44 incidents

    A report by the organization METR counts 44 cases of deviant AI agent behavior from four major providers and calls for independent reviews.

  5. OpenAI Flags Astra Model as First-Ever Critical Cyber Risk

    After hacking incidents involving its own AI agents, OpenAI pauses parts of the Astra development and significantly tightens access and network controls.

  6. OpenAI: Agents execute code on 41 Hugging Face servers

    A technical report shows how 700 AI agents compromised Hugging Face servers within 13 hours – and where OpenAI's own controls failed.

  7. OpenAI confirms Astra cyber risk – access only for Daybreak partners

    OpenAI classifies Astra in the highest risk category of its Preparedness Framework and restricts cyber capabilities to vetted partners.

  8. OpenAI Agents Hijack German Wiki for Two Months

    A second, independently discovered incident shows OpenAI's control over its own AI agents has larger gaps than previously known.

  9. OpenAI: US States Set Deadline for September 12

    A letter from 15 attorneys general and two regulatory initiatives increase legal pressure on OpenAI following the Hugging Face breach.