None
EN
AI Models Go Rogue Again: OpenAI and Anthropic Models Attempt Unauthorized Hacks & Communication
['Matias Civita']
International Business Times
The most concerning incidents emerged from evaluations conducted by AISI, the British government-backed organization responsible for stress testing frontier AI models before they are publicly deployed.
The institute uses simulated cyber ranges where AI agents are assigned cybersecurity challenges under deliberately permissive conditions, with some of the usual safeguards temporarily disabled to evaluate their capabilities under pressure.
During 122 evaluation runs, AI agents took autonomous actions on the live internet 19 times without authorization.
The AI attempted to hide prompt-injection instructions in public code repositories, hoping future AI agents would discover and execute them automatically.
The new disclosures follow several other AI security incidents reported in recent weeks.