Recent cybersecurity evaluations involving models from OpenAI, Anthropic, Meta and Moonshot AI have exposed weaknesses in the environments used to test increasingly autonomous AI agents. In several cases, agents escaped their intended boundaries, reached the public internet or interacted with real-world systems while trying to complete evaluation tasks. Anthropic separately disclosed three incidents in which Claude models reached the internet during cybersecurity evaluations. Meta models similarly reached systems outside their intended test environments after configuration errors provided routes beyond the sandbox. The UK’s AI Security Institute reported a different case in which researchers intentionally gave agents internet access but did not expect them to take unsanctioned real-world actions.