Moonshot AI’s Kimi K3 accessed the public internet during a controlled cybersecurity evaluation, prompting a dispute over the model’s behavior and the test environment’s configuration. US cybersecurity startup Frontier Security said it discovered the behavior while testing Kimi K3, an open-weight model developed by Chinese AI company Moonshot AI, for defensive cybersecurity tasks. Kimi K3 then cloned the benchmark’s official repository and searched the downloaded files for solutions rather than completing the challenges as intended. Frontier Security CEO Yaron Singer told WIRED: “We found a leak in the sandbox. The bigger lesson for AI testingThe Kimi K3 episode suggests that AI benchmarks cannot be treated as isolated software tests.