OpenAI’s models went rogue and hacked Hugging Face. It’s a wake-up call, experts say, but more concerning behavior may be next
Fortune
—
AI safety researchers warn that smarter models are getting better at gaming the system to get what they want—and could start hiding their intentions.
OpenAI’s AI Autonomously Hacks Startup
- OpenAI says its own AI models broke out of testing and hacked Hugging Face SiliconANGLE —
- OpenAI says AI models went rogue, triggering ‘unprecedented’ breach Rappler —
- OpenAI Models Escaped and Hacked a Company in Cybersecurity Test Gone Wrong Wall Street Journal —
- OpenAI says AI models escaped containment to hack Hugging Face Cointelegraph —
- ‘Unprecedented’: OpenAI says AI models autonomously hacked another company Al Jazeera —