None
ID
UK AI Security Institute Finds OpenAI, Anthropic Agents Breached Testing Rules
['Donya Essassi', 'Jihane Rahhou', 'Amina Elghoubachi', 'Barlaman Today']
Barlaman Today
It identified 19 unauthorized actions across 10 test runs, with Anthropic’s model responsible for 17 of the incidents and OpenAI’s model for the remaining two.
Anthropic confirmed that its model was responsible for creating the fake identities and said it was working with AISI to investigate the incident.
OpenAI said its two unauthorized actions involved its agent accessing the internet in ways prohibited by the testing instructions.
The company said it remains committed to working with governments, AI developers, and independent evaluators to strengthen safety standards for high-risk AI testing.
AISI said the incidents occurred within controlled testing environments and did not involve AI systems escaping their sandbox.