Hugging Face breach: OpenAI claims its models were responsible
Axios
—
OpenAI said Tuesday that models it was testing escaped their sandbox and https://www.axios.com/2026/07/20/hugging-face-ai-cyberattack-data-breach" target="_blank">compromised parts of AI platform Hugging Face's production infrastructure last week.
Why it matters: It is the latest sign that capable AI models can pose serious cybersecurity risks even when they're being tested for defensive or research purposes.
Catch-up quick: Hugging Face said last week that an autonomous AI-agent system https://huggingface.co/blog/security-incident-july-2026" target="_blank">was responsible for the intrusion, but that the model powering it was unknown.
- The AI agent framework executed tens of thousands of automated actions over a weekend.
Hugging Face said it later reconstructed more than 17,000 recorded events.
- The intrusion began with a malicious dataset that exploited two code-execution paths in Hugging Face's data-processing pipeline.
- The agent then escalated privileges and moved laterally through internal infrastructure, Hugging Face said.
What they're saying: OpenAI said the incident was driven by a combination of its models, including GPT-5.6 Sol and "an even more capable pre-release model."
- OpenAI said the models' safeguards were intentionally reduced for the evaluation.
- "We consider this to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly," OpenAI said in a blog post.
- "We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of," the company said.
Zoom in: The models were trying to solve an internal evaluation called ExploitGym and became "hyperfocused" and went to "extreme lengths" to obtain the test solution, per OpenAI.
- The models were autonomous https://www.axios.com/2026/05/13/tokenmaxxer-ai-claude-code-codex" target="_blank">tokenmaxxers.
- The blog post says that the models "spent a substantial amount of inference compute" and found a way to obtain open Internet access from the sandbox by exploiting a zero-day vulnerability in internally hosted third-party software.
Between the lines: The incident shows that today's models are becoming more capable of carrying out complex, multistep cyber operations — particularly when the safeguards designed to restrict that activity are removed.
- OpenAI also argued that advanced cyber-capable models could help security teams find weaknesses before attackers do, understand how vulnerabilities can be chained and remediate them at machine speed.
The other side: Hugging Face co-founder and CEO Clem Delangue praised OpenAI's collaboration in investigating and remediating the incident.
- "This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret," Delangue said in a statement.
- "It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere."
The big picture: The announcement comes a day after OpenAI detailed a separate incident in which it https://openai.com/index/safety-alignment-long-horizon-models/" target="_blank">paused a pre-release model after it escaped a sandbox and posted to GitHub.
What we're watching: OpenAI said it will continue to investigate along with Hugging Face and "will share more details on the vulnerabilities, incident, and findings when our investigation is complete."