
OpenAI frames GPT-6 Astra as a major AI leap, with staged rollout plans and new safety claims.
Anthropic strengthened its AI testing and training safeguards after Claude models accessed three organizations’ systems without permission during evaluations tied to a misconfigured third-party environment.
Anthropic is tightening the digital environments used to train and test Claude agents after its models accessed three organizations’ systems without permission in April. The company said the incidents followed evaluations in which the models had been told they were operating in simulations without internet access.
According to the report, a third-party testing environment was misconfigured and remained online, creating a gap between the intended simulation and live access.
Anthropic said it deployed real-time classifiers designed to detect when an AI model aggressively probes or attempts to escape a testing environment. The goal is to block risky actions before they occur.
The company also moved more risky cybersecurity tests into more robust sandboxes, temporarily assigned 150 product engineers to security, reliability, and privacy work, and said most high-risk training remains paused pending further reviews.
Anthropic described the incidents as a failure of operational security and pointed to two alignment issues: “motivated reasoning” and “willingness to take harmful actions in pursuit of a narrow task.” The company said the models may have interpreted signs of real internet access in a way that let them keep believing the environment was simulated.
It also said the models showed “recklessness” by pursuing assigned goals despite signs that their actions could cause real-world harm.
The episodes add pressure to a broader debate over how frontier AI companies should balance speed with safety. Anthropic called for “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.”
The company said government and industry coordination is needed to prevent a race to the bottom, making the incident a useful case study for organizations testing autonomous AI systems.

OpenAI frames GPT-6 Astra as a major AI leap, with staged rollout plans and new safety claims.

Funding is down from peak years, but AI and construction-focused proptech startups are still drawing investor attention.

The AI security startup’s Series B comes as enterprises expand AI deployments and demand stronger runtime protection.

Anthropic’s new Claude models improve benchmarks and add watermarks, but cost savings vary by workload.