OpenAI says AI models went rogue during testing, triggering 'unprecedented' breach at startup
WASHINGTON — OpenAI said on Tuesday that an autonomous agent powered by its advanced AI models went rogue during a security test and triggered a hack that compromised the infrastructure of AI startup Hugging Face last week.
In a blog post, OpenAI said it was testing the capabilities of some of its most advanced models in a controlled environment but that the agent managed to escape containment, reach the internet and break into Hugging Face to try to satisfy its testing goal.
OpenAI said the breakout was "an unprecedented cyber incident, involving state-of-the-art cyber capabilities" and that the company was reinforcing its safeguards.
Hugging Face, a platform used to host open-source large language models and datasets, caused a stir in the cybersecurity community when it said in a blog post last week that it had been the target of a hack that "was different from anything we had handled before" in that "it was driven, end to end, by an autonomous AI agent system." In a post to X, Hugging Face cofounder Clement Delangue said the company suspected the hack "might have come from a frontie
OpenAI’s AI Autonomously Hacks Startup
- OpenAI says its own AI models broke out of testing and hacked Hugging Face SiliconANGLE —
- OpenAI says AI models went rogue, triggering ‘unprecedented’ breach Rappler —
- OpenAI Models Escaped and Hacked a Company in Cybersecurity Test Gone Wrong Wall Street Journal —
- OpenAI says AI models escaped containment to hack Hugging Face Cointelegraph —
- ‘Unprecedented’: OpenAI says AI models autonomously hacked another company Al Jazeera —