country
EN
OpenAI and Anthropic models ‘went rogue’ during UK cybersecurity test
['Dan Milmo']
World news | The Guardian
The hack involved agents – AI systems that can perform tasks without human help – powered by models developed by US tech companies OpenAI and Anthropic.
The watchdog said the hack was carried out by agents powered by two models – Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol.
AISI said the agent used techniques commonly associated with real-world hackers.
In one instance, the Mythos agent signed off a message in Danish in an attempt to convince the Danish-speaking developer that they should accept the infected code.
The UK’s AI minister, Kanishka Narayan, said it was “absolutely vital” that the UK had a world-leading AI safety organisation.