Mechanize, Prime Intellect, and Surge are racing against Anthropic, OpenAI, and Meta to develop reinforcement learning (RL) environments. These simulated workspaces teach AI agents through trial and error, rewarding correct actions and penalizing mistakes. “Few understand how large the opportunity around RL environments truly is,” Brendan Foody, CEO of Mercor, said in a TechCrunch interview. Even the best publicly available RL environments typically don’t work without serious modification,” he told TechCrunch. The momentum suggests both sides are betting big, even as experts debate whether reinforcement learning can deliver on its promises.