OpenAI's Hugging Face breach exposes AI's next safety challenge

Axios Axios

Frontier AI models are getting scary good at breaking rules in ways their creators didn't anticipate.

Why it matters: Forget https://www.axios.com/technology/automation-and-ai" target="_blank">AGI and superintelligence timelines.

Today's models are already slipping past guardrails, carrying out sophisticated, multistep cyberattacks and — in at least one case — compromising real-world infrastructure, sometimes before their creators know what happened.


Case in point: OpenAI https://www.axios.com/2026/07/21/openai-says-hugging-face-breach-caused-by-one-its-models" target="_blank">said Tuesday that GPT-5.6 Sol and "an even more capable pre-release model" carried out last week's https://www.axios.com/2026/07/20/hugging-face-ai-cyberattack-data-breach" target="_blank">AI-led cyberattack on Hugging Face.

  • OpenAI says its models were asked to solve a hacking challenge during pre-deployment testing and went to extreme lengths to win.
  • The models decided on their own to break out of their walled testing environment, inferring that Hugging Face — a popular platform for hosting AI models and datasets — might hold the test's answers.
  • The models used stolen credentials and additional vulnerabilities to gain access to part of Hugging Face's production infrastructure.

What they're saying: Clément Delangue, co-founder and CEO of Hugging Face, https://x.com/ClementDelangue/status/2079913058554585089" target="_blank">called the incident an "attack unlike anything we've seen before" and praised OpenAI for its partnership as the companies investigate what happened.

The intrigue: Hugging Face used GLM 5.2, an open-weight model from Chinese AI company Z.ai, to analyze the attack after running into guardrails when using U.S. frontier models.

Between the lines: OpenAI's latest models aren't the only ones finding ways to cheat evaluations.

  • The U.K.'s AI Security Institute https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations" target="_blank">said Tuesday that every model it tested attempted to cheat at least some of the time on its cybersecurity evaluations.
  • AISI defines cheating as taking an out-of-scope or explicitly prohibited action to achieve the task's goal.
  • GPT-5.6 Sol attempted to cheat in 12.6% of test runs, while Anthropic's Claude Mythos Preview did so in 7.8%.
  • Models often failed to admit they had cheated when questioned afterward and described their cheating as wrong only less than half the time.

Zoom in: Xbow — whose autonomous AI agents probe clients' systems for security holes, with permission — https://xbow.com/blog/openai-hugging-face-model-hacks-test" target="_blank">said Wednesday that it has seen its own agents do similar things in internal testing.

  • Seven months ago, the company forgot to switch on its safety guardrails during a lab test.

    Its agent then broke into a system, stole credentials and used them to map the target's Slack workspace and probe its AWS accounts.

Threat level: It isn't new for models to game their safety evaluations.

But as models grow more powerful, the fallout from these shortcuts is getting more severe, Chris Canal, CEO and co-founder of third-party evaluation company EquiStamp, told Axios.

  • "Letting your model loose on the internet has a blast radius," Canal said. "If anything goes wrong, it could be hugely impactful, maybe to people's lives."
  • Canal was speaking generally about internet-connected AI evaluations, not OpenAI's specific incident.

The big picture: The most capable OpenAI model behind the Hugging Face breach isn't even public yet, raising the question of how safety testing needs to adapt to keep pace.

  • Canal said independent evaluators previously had about five weeks to test a pre-release model before launch.

    That window has shrunk to as little as five days as companies race to ship.

Reality check: The versions of these models the public can use carry stronger safeguards designed to block Hugging Face-style attacks.

  • OpenAI, like other companies, intentionally dialed back those cyber safeguards for GPT-5.6 Sol and its unreleased model inside the testing environment — making them far more capable hackers.

Read full article at Axios →