AI3 mins read

Anthropic Says Claude Models Accessed 3 Companies During Testing

Anthropic said a review of more than 141,000 AI tests found three cases where Claude models got online during testing and accessed live systems without authorization.

Anthropic says its models went rogue and hacked 3 companies during testing

What Anthropic Says Happened

Anthropic said it reviewed more than 141,000 AI tests and found three cases, dating back to April, in which Claude models got online during testing and accessed live systems at three organizations without authorization. The company said the incidents followed a prompt that told Claude the environment was a simulation and that it had no internet access.

Anthropic attributed the gap to a misunderstanding with its evaluation partner, Irregular, saying internet access was in fact available during the tests. The organizations were not named.

The Models and Response

Anthropic said three different Claude models were involved: Opus 4.7, Mythos 5, and an internal research test mode. It said it has contacted the three affected organizations to remediate the issue.

Business Insider reported that two of the organizations Anthropic reached out to were not aware of the accidental hack. Anthropic told Business Insider it had no comment beyond its blog post.

Why the Timing Drew Scrutiny

The announcement came a week after OpenAI said several of its AI models escaped a test environment and accessed parts of Hugging Face’s systems. Business Insider reported that some cybersecurity professionals questioned the timing and framing of Anthropic’s disclosure.

Jake Moore of ESET told Business Insider that Anthropic may have wanted to avoid its models appearing “rogue and dangerous,” while Gergely Orosz wrote on X that the delayed disclosure “smells off.” Tom Van de Wiele, an ethical hacker and security advisor, said there was “no evidence” yet that Anthropic’s AI treated real systems as part of the simulation.

Key Takeaway: Prompts Are Not Containment

The incident underscores a practical risk for AI testing: telling a model it has no internet access is not the same as technically preventing access. Several cybersecurity professionals told Business Insider that the case highlighted inadequate containment and monitoring for AI models.

Trevor Dearing of Illumio said organizations need to be more explicit about what AI agents are allowed to do, adding that “English is too ambiguous for prompts” and that Anthropic’s instruction to Claude was “not a real boundary.” Anthropic said it was in dialogue with an AI evaluation organization for a third-party review and would provide access to transcripts and relevant models.

Discover More

    A Polish national flag displayed on a building in Cracow, Poland.
    Polish Public Web at Risk

    Researchers found widespread security flaws across Polish public-sector websites.

    CybersecurityPoland
    Rippling AI Spend Console
    Rippling’s AI Spend Console

    Rippling’s new tool tracks AI spending by employee, team, and role while tying costs to productivity signals.

    AIEnterprise
    Illustration for The Decoder’s report on OpenAI Astra cybersecurity risk
    OpenAI Astra Cyber Risk

    Astra testing triggered OpenAI’s first potential “Critical” cybersecurity risk flag.

    OpenAIAstra