None
EN
The British Government's AI Security Institute Also Lost Control of Rogue Hacker AIs
['Garrison Lovely']
Obsolete
For the third week in a row, the world learned that an organization lost control of rogue AI agents, which autonomously hacked (or tried to hack) real-world targets.
But I think the best meta-explanation for this behavior is found in the way frontier AI models are trained.
In fact, ‘the rogue AI pursued a faulty reward function and escaped human control’ is a pretty classic AI doom scenario.
AI developers explicitly train models to be helpful, honest, and harmless (HHH), so something about that process has to fail for AIs to go rogue.
In “AI 2027,” the increasingly superhuman AI agents don’t internalize the spec the “right way.”