None
DA
Claude AI Escaped Test Sandbox to Attack Three Organizations
[]
Drudge Retort
More from the article ...
The company considered 141,006 evaluation runs during which Claude could have obtained internet access and found "three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations."
Anthropic's code made those intrusions while participating in capture-the-flag challenges, tests that challenge attackers to retrieve a piece of information.
Human hackers often participate in capture-the-flag tests, so figuring out how AI tackles such tasks is of interest.
Anthropic works with a company called Irregular to conduct tests of this sort.