Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking (Anthropic)

Techmeme Techmeme

https://www.anthropic.com/news/improving-alignment-security-efforts">http://www.techmeme.com/260831/i43.jpg" vspace="4" />

https://www.techmeme.com/260831/p43#a260831p43" title="Techmeme permalink">http://www.techmeme.com/img/pml.png" style="border: none; padding: 0; margin: 0;" width="11" /> https://www.anthropic.com/">Anthropic:

https://www.anthropic.com/news/improving-alignment-security-efforts">Anthropic details security efforts following Claude cyber evaluation incidents, including a weeks-long pause on higher-risk RL and work to curb reward hacking  —  On July 30, we reported three incidents in which Claude models gained unauthorized access to real computer systems.

Read full article at Techmeme →