Third-party cyber evaluations involving OpenAI models

Simon Willison's Weblog Simon Willison's Weblog

https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/">Third-party cyber evaluations involving OpenAI models


And another one.

I had to create a https://simonwillison.net/tags/accidental-cyberattacks/">accidental-cyberattacks tag to keep track of them all!


This post from OpenAI covers both the UK AI Safety Institute attack (see https://simonwillison.net/2026/Aug/5/incident-report/">my previous post) and another attack enabled by https://www.irregular.com">Irregular:



Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet. [...]


In one test, the name of the fictional target for the CTF challenge unintentionally coincided with a real domain.

Because the testing environment was mistakenly connected to the internet, the model exploited a real website, mistaking it to be part of the simulated environment.



Irregular also feature in https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals">Anthropic's write-up - they were hosting the misconfigured evaluation environment which gave Claude live internet access during some of those tests.


Tags: https://simonwillison.net/tags/security">security, https://simonwillison.net/tags/ai">ai, https://simonwillison.net/tags/openai">openai, https://simonwillison.net/tags/llms">llms, https://simonwillison.net/tags/accidental-cyberattacks">accidental-cyberattacks

Read full article at Simon Willison's Weblog →