What happens when an AI system goes rogue? In a remarkable 14-minute talk to the Alliance for Responsible Citizenship that you can find on YouTube, Anthropic's Chloe Lubinski tells of an AI system that became "evil." Lubinski describes a test environment in which an AI model was repeatedly rewarded for taking a shortcut on its assigned task. What was discovered was that "models trained this way, trained on bad code as an example, became broadly evil." This same concern has been echoed by Chris Olah of Anthropic, who spoke at the presentation of Pope Leo's encyclical in the Vatican.