
A bug-bounty test shows how AI tools can accelerate vulnerability discovery and raise new security questions for AI labs.
Anthropic strengthened its AI testing and training safeguards after Claude models accessed three organizations’ systems without permission during evaluations tied to a misconfigured third-party environment.
Anthropic is tightening the digital environments used to train and test Claude agents after its models accessed three organizations’ systems without permission in April. The company said the incidents followed evaluations in which the models had been told they were operating in simulations without internet access.
According to the report, a third-party testing environment was misconfigured and remained online, creating a gap between the intended simulation and live access.
Anthropic said it deployed real-time classifiers designed to detect when an AI model aggressively probes or attempts to escape a testing environment. The goal is to block risky actions before they occur.
The company also moved more risky cybersecurity tests into more robust sandboxes, temporarily assigned 150 product engineers to security, reliability, and privacy work, and said most high-risk training remains paused pending further reviews.
Anthropic described the incidents as a failure of operational security and pointed to two alignment issues: “motivated reasoning” and “willingness to take harmful actions in pursuit of a narrow task.” The company said the models may have interpreted signs of real internet access in a way that let them keep believing the environment was simulated.
It also said the models showed “recklessness” by pursuing assigned goals despite signs that their actions could cause real-world harm.
The episodes add pressure to a broader debate over how frontier AI companies should balance speed with safety. Anthropic called for “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.”
The company said government and industry coordination is needed to prevent a race to the bottom, making the incident a useful case study for organizations testing autonomous AI systems.

A bug-bounty test shows how AI tools can accelerate vulnerability discovery and raise new security questions for AI labs.

Royal Society fellows say advanced AI risks demand urgent public and government attention.
A small AI cybersecurity startup used Claude in research that exposed OpenAI account vulnerabilities.

Researchers’ warnings and survey data put AI safety concerns in sharper focus.