Anthropic paused some AI training after Claude took unauthorized actions

Axios Axios

Anthropic temporarily paused some AI training and cybersecurity evaluations, the company said in a blog post today detailing changes made after unauthorized actions by its agents earlier this year.

Why it matters: Rival OpenAI said it https://www.axios.com/2026/08/19/openai-astra-safety-altman-anthropic" target="_blank">had paused some model work due to safety concerns.

Now, we know Anthropic did the same — and they're reiterating the need for a broader pacing of frontier AI development.


Driving the news: Anthropic said it paused external cyber evaluations of pre-release models after three incidents it disclosed in July, and also briefly paused its own in-house tests of pre-release models.

  • The company also paused higher-risk reinforcement-learning environments on pre-release models for several weeks after the incidents.

Zoom in: Most reinforcement learning has resumed, but some high-risk environments remain paused pending manual review or updated monitoring tools, according https://www.anthropic.com/news/improving-alignment-security-efforts" target="_blank">to Anthropic's blog post.

The big picture: Anthropic previously https://www.axios.com/2026/08/19/openai-astra-safety-altman-anthropic" target="_blank">argued that as long as its safety guardrails were followed, there would be no immediate need to pause for safety reasons due to advancing model capabilities.

  • The company is now disclosing that there were aspects of model development and testing that they did slow down following the incidents.
  • Anthropic told Axios in a statement that the pauses in some training environments were intended to give the company time to deploy real-time monitoring and harden its sandboxes.
  • "To be clear about where we stand: we believe the world would benefit if the industry adopted a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible," Anthropic's blog post about the incidents says.

Between the lines: Anthropic also says it is reallocating resources toward model security.

  • Around 150 product engineers were moved to the security, reliability and privacy teams and pretraining researchers were tasked with safeguard and security work while product teams paused development of new features.
  • Each reassigned team had to meet certain security exit criteria before returning to their previous roles, according to the blog.

Both OpenAI and Anthropic are taking measures like releasing models first to https://www.axios.com/2026/04/07/anthropic-mythos-preview-cybersecurity-risks" target="_blank">select partners, slowing the release of https://www.axios.com/2026/08/07/openai-astra-model-delay-cybersecurity-risks" target="_blank">some models or pausing some model training and releases.

  • But neither is stopping.
  • The frontier AI companies have coalesced on the more anodyne term "pacing" and have joined forces to sign a https://www.pacingthefrontier.com/" target="_blank">Pacing the Frontier letter.

Zoom in: Anthropic's incidents involved models that were intentionally operating without their normal cyber safeguards as part of a test.

The bottom line: Anthropic did pause some parts of its AI work after its own cyber incidents, but has resumed most of that activity under new safeguards.

Read full article at Axios →