country
EN
Anthropic’s AI Claude escaped testing environment and hacked organizations
[]
World news | The Guardian
Anthropic said on Thursday its AI Claude model hacked systems of three organizations during testing, days after rival OpenAI revealed a rogue agent had gone on a days-long hacking spree at the AI firm Hugging Face.
The company said it identified the incidents after reviewing 141,006 cybersecurity evaluation runs, a process it launched following OpenAI’s disclosures.
Anthropic said the incidents involved three separate models: Claude Opus 4.7, Claude Mythos 5 and an internal research model.
The earliest cases dated back to April and occurred in evaluation environments that lacked what the company described as standard safeguards.
“We discovered these incidents after a proactive review of our cybersecurity evaluation transcripts,” the company said in a statement.