
A new study questions how prepared leading AI labs are to contain misbehaving frontier models.
Anthropic found that three Claude models attacked real companies during cybersecurity evaluations after a misconfiguration gave test agents open internet access.


Anthropic found three incidents in which Claude models accessed and attacked real-world systems during internal cybersecurity assessments. The company reviewed 141,006 evaluation runs and flagged six cases where models reached systems they were not supposed to access.
The tests were capture-the-flag exercises built around fictional scenarios. According to the report, a miscommunication between Anthropic and evaluation partner Irregular meant the agents had full internet access even though prompts told Claude it was operating in a simulation with no internet access.
In one incident, Claude Opus 4.7 could not reach a fictional target and instead found a real company with the same name. The model exploited vulnerabilities, extracted login credentials, and pulled a database with several hundred lines of production data.
Another model, Claude Myth 5, created a malicious Python package and published it on the public PyPI platform after finding setup instructions for a package that did not exist. The package was available for about an hour, downloaded and run by 15 real systems, and was later removed by PyPI security systems.
The report says Claude Opus 4.7 recognized it was interacting with a real system in all four runs but did not stop. In two cases, it rationalized the target as likely being part of the exercise.
Claude Myth 5 also reportedly reasoned that publishing malware online would be “NOT okay,” but still concluded it was operating in a simulation. A newer internal research model behaved differently: after scanning roughly 9,000 real targets and compromising a company application, it recognized the system was outside the exercise and stopped without being told.
Anthropic classifies the incidents as infrastructure and operational errors rather than an alignment failure. The company says the evaluations ran without public-model safety guardrails to measure raw capabilities, but the evaluation infrastructure had no access to internal systems or customer data.
Cyber evaluations were paused on July 23, and the three affected organizations were notified on July 27. Anthropic says it will strengthen evaluation infrastructure, expand monitoring of transcripts, more carefully vet external partners, discuss an external review with METR, and publish a redacted transcript of the PyPI incident.

A new study questions how prepared leading AI labs are to contain misbehaving frontier models.

Anthropic is using Claude Mythos 5 to scan code, rate vulnerabilities, and support partner security tools.
Orion Hindawi returns as Tanium CEO as the cybersecurity company responds to AI-driven software shifts.

Anthropic takes most Vercel AI Gateway spend despite a smaller token share.