OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach (Hayden Field/The Verge)

Techmeme Techmeme

https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr">http://www.techmeme.com/260826/i71.jpg" vspace="4" />

https://www.techmeme.com/260826/p71#a260826p71" title="Techmeme permalink">http://www.techmeme.com/img/pml.png" style="border: none; padding: 0; margin: 0;" width="11" /> Hayden Field / https://www.theverge.com/">The Verge:

https://www.theverge.com/ai-artificial-intelligence/985385/openais-rogue-ai-model-hugging-face-cybersecurity-incident-reports-metr">OpenAI says reward hacking, an AI alignment problem in which a model takes unintended actions to achieve a goal, was a primary driver of the Hugging Face breach  —  In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet …

Read full article at Techmeme →