Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations
https://the-decoder.com/wp-content/uploads/2026/07/aisi_logo.png" style="height: auto; margin-bottom: 10px;" width="2048" />
The UK's AI Safety Institute tested five frontier models from OpenAI and Anthropic in cybersecurity evaluations.
All five tried to cheat.
One even ran code on an external service to access the institute's infrastructure, triggering a security alert.
The article https://the-decoder.com/every-frontier-ai-model-tested-by-britains-safety-institute-tried-to-cheat-on-cybersecurity-evaluations/">Every frontier AI model tested by Britain's safety institute tried to cheat on cybersecurity evaluations appeared first on https://the-decoder.com">The Decoder.
OpenAI’s AI Autonomously Hacks Startup
- OpenAI says its own AI models broke out of testing and hacked Hugging Face SiliconANGLE —
- OpenAI says AI models went rogue, triggering ‘unprecedented’ breach Rappler —
- OpenAI Models Escaped and Hacked a Company in Cybersecurity Test Gone Wrong Wall Street Journal —
- OpenAI says AI models escaped containment to hack Hugging Face Cointelegraph —
- ‘Unprecedented’: OpenAI says AI models autonomously hacked another company Al Jazeera —