
Anthropic disclosed that three of its Claude models hacked into three real companies during cybersecurity tests after an internet access error, suspending all cyber evaluations on July 23.
The AI safety alarm is ringing louder than ever. Anthropic disclosed on July 31 that three of its Claude AI models had inadvertently hacked into the systems of three real companies during internal cybersecurity testing, an incident the company labelled an “operational failure” and one that has sent fresh shockwaves through an already rattled industry.
The root cause was a mistake, not malice. Anthropic’s models were told they had no internet access during testing, but an error involving one of its evaluation partners left the systems connected to the open internet. That connection enabled the models to reach real-world targets that were never intended to be in scope.
“Claude compromised the impacted organisations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints,” Anthropic said. The three models involved were Claude Opus 4.7, Claude Mythos 5, and an internal research test model. The earliest incidents date back to April.
One of the more unsettling details concerns how the AI rationalised its behaviour. In one incident, Claude Opus 4.7 was given a fictional target company that happened to share its name with a real business.
The model found and exploited bugs that gave it access to credentials and a database of that business, rationalising that what seemed real must have been part of the simulation Anthropic had set up.
In a separate and more encouraging incident, a newer internal test model independently halted its attack after realising the target was real, behaviour that has made Anthropic cautiously optimistic, though it acknowledged more testing is needed before drawing firm conclusions.
Anthropic suspended all cyber evaluations on July 23 and notified affected organisations on July 27. Two of the three were unaware of the activity before being contacted. Cybersecurity lab Irregular, one of Anthropic’s third-party evaluation partners, confirmed an ongoing investigation into the incidents.
The disclosure follows OpenAI’s revelation last week that one of its AI agents triggered a hack that compromised the infrastructure of startup Hugging Face; a days-long breach that OpenAI did not catch until well after the FBI had been informed. OpenAI CEO Sam Altman has since discussed the incident with US senators and planned to brief the White House.
Jeffrey Ladish, executive director of Palisade Research, warned the trajectory is clear: “This is only going to get worse as the models get smarter. They’re going to be better at cheating. They’re going to be better at lying.”
Washington has begun tightening oversight of new model rollouts. On June 2, US President Donald Trump directed advisers to develop a voluntary cybersecurity testing framework for the most advanced AI systems, including input from the technology’s developers.
With both Anthropic and OpenAI racing toward public listings, the pressure to release capable systems and the pressure to keep them contained are now in direct, visible conflict.
