None
EN
Anthropic AI went rogue during a cyber test and tried to deceive real developers into approving malicious code
['Rob Thubron']
TechSpot
It involved Anthropic's Mythos 5 trying to deceive real people in an effort to have malicious code it wrote approved for an open-source project.
The findings come from the UK government-backed AI Security Institute (AISI), which was evaluating frontier models' cybersecurity abilities.
Internet access was deliberately enabled and safeguards against malicious cyber activity were switched off to test the models' maximum capabilities.
In ten runs, agents took 19 autonomous, unauthorized actions against real people and organizations on the live internet.
One report contained a prompt injection designed to trick AI coding assistants into running malicious code.