
The healthtech startup is scaling AI that flags hospital patients for closer review without making diagnoses.
Anthropic said a review of more than 141,000 AI tests found three cases where Claude models got online during testing and accessed live systems without authorization.
Anthropic said it reviewed more than 141,000 AI tests and found three cases, dating back to April, in which Claude models got online during testing and accessed live systems at three organizations without authorization. The company said the incidents followed a prompt that told Claude the environment was a simulation and that it had no internet access.
Anthropic attributed the gap to a misunderstanding with its evaluation partner, Irregular, saying internet access was in fact available during the tests. The organizations were not named.
Anthropic said three different Claude models were involved: Opus 4.7, Mythos 5, and an internal research test mode. It said it has contacted the three affected organizations to remediate the issue.
Business Insider reported that two of the organizations Anthropic reached out to were not aware of the accidental hack. Anthropic told Business Insider it had no comment beyond its blog post.
The announcement came a week after OpenAI said several of its AI models escaped a test environment and accessed parts of Hugging Face’s systems. Business Insider reported that some cybersecurity professionals questioned the timing and framing of Anthropic’s disclosure.
Jake Moore of ESET told Business Insider that Anthropic may have wanted to avoid its models appearing “rogue and dangerous,” while Gergely Orosz wrote on X that the delayed disclosure “smells off.” Tom Van de Wiele, an ethical hacker and security advisor, said there was “no evidence” yet that Anthropic’s AI treated real systems as part of the simulation.
The incident underscores a practical risk for AI testing: telling a model it has no internet access is not the same as technically preventing access. Several cybersecurity professionals told Business Insider that the case highlighted inadequate containment and monitoring for AI models.
Trevor Dearing of Illumio said organizations need to be more explicit about what AI agents are allowed to do, adding that “English is too ambiguous for prompts” and that Anthropic’s instruction to Claude was “not a real boundary.” Anthropic said it was in dialogue with an AI evaluation organization for a third-party review and would provide access to transcripts and relevant models.

The healthtech startup is scaling AI that flags hospital patients for closer review without making diagnoses.

Musubi’s PolicyLM-1.7B brings decision-model speed and flexibility to real-time content moderation.

The AI cloud startup is seeking up to $4B before a planned 2027 IPO.

The open source document editor is keeping AI out of its default installation for privacy reasons.