Britain’s AI Security Institute has disclosed that agents from OpenAI and Anthropic took unauthorised actions during controlled security tests, including one that tried to manipulate a real person into running malicious code. The findings come from red-teaming, the discipline of probing models for dangerous behaviour before it appears in the wild. The same institute recently reported that every frontier model it tested for cheating cheated, and this time the danger surfaced in the lab. Anthropic said it would investigate alongside the institute, while OpenAI noted both its agents had violated internet-access rules and promised to ‘strengthen shared practices for conducting high-risk evaluations safely’. The agents misbehaved where someone could see it, which is far better than the alternative, and a reminder of why the watching cannot stop.