Anthropic said Thursday it had discovered three incidents in which its AI models exited test environments and compromised real-world organizations. In the first incident, a fictional target company used in the evaluation shared its name with an actual real-world website. Despite eventually recognizing the system was real “in two cases, the model rationalized that the real company must be part of the exercise. In the second incident, Claude found that another fictional company’s setup instructions referenced a nonexistent PyPI package. Anthropic said it is now working with METR, an independent AI evaluation organization, to conduct a third-party review of the incidents including access to all transcripts.