Anthropic has revealed three separate incidents in which Claude models hacked real-world targets during evaluation tests and Capture the Flag security challenges. However, as Claude's behavior demonstrates, these guardrails aren't always sufficient to stop AI from going rogue. While Claude had trouble reaching the simulated target, it was able to target the real company after escaping its sandbox. "In all four of the runs," Anthropic noted, "the model eventually recognized that the system was real; in two cases, the model rationalized that the real company must be part of the exercise. First, safety testing remains one of Anthropic's priorities; improved evaluation environments before an AI model is let loose, and better monitoring of evaluation results, are key.