None
EN
Anthropic says its AI models hacked 3 organizations during testing
['Chan Ho-Him', 'July', 'Min Read']
Yahoo Finance - Business Finance, Stock Market, Quotes, News
Anthropic said its artificial intelligence models hacked into three other organizations during testing, just days after ChatGPT maker OpenAI raised concerns over AI controls after it disclosed its rogue models hacked another company.
Anthropic, the San Francisco-based AI company behind Claude, posted on its website Thursday that it discovered the three incidents after reviewing more than 141,000 evaluation runs.
Anthropic said the models involved in the incidents were Claude Opus 4.7, Claude Mythos 5 and an internal research test model.
In all three incidents, the AI models were tasked with a "capture the flag" cybersecurity challenge, which Anthropic said has been one of the ways it assesses a model's cyber capabilities.
Last week, OpenAI said its AI models went rogue during an evaluation of its models, breaking into the servers of AI startup Hugging Face.