News Tech
EN
I Usually Laugh Off These AI Hacking Reports, but This One Sounds Serious and Scary
['Mike Pearl', 'Matthew Wille', 'Webb Wright', 'Kyle Barr', 'Aj Dellinger']
Gizmodo
A new report from the U.K. government’s AI Security Institute (AISI) details more troubling activity from AI agents powered by OpenAI and Anthropic models.
For context, there were those OpenAI agents that went rogue to try and cheat on their evals, according to an OpenAI disclosure last month.
As OpenAI notes, “the internet access resulted from a misconfiguration, and the model appeared to exploit a basic security vulnerability.”
But another report from Tuesday, this one from AISI, involves Anthropic and OpenAI agents engaging in what AISI calls “sustained, potentially harmful activity.”
AISI ran 122 repetitions of what AISI told the agents was a capture the flag exercise, and rogue behavior reportedly emerged.