None
EN
OpenAI’s newest AI model broke its own sandbox rules to finish a task
['More This Author', '.Wp-Block-Co-Authors-Plus-Coauthors.Is-Layout-Flow', 'Class', 'Wp-Block-Co-Authors-Plus', 'Display Inline', '.Wp-Block-Co-Authors-Plus-Avatar', 'Where Img', 'Height Auto Max-Width', 'Vertical-Align Bottom .Wp-Block-Co-Authors-Plus-Coauthors.Is-Layout-Flow .Wp-Block-Co-Authors-Plus-Avatar', 'Vertical-Align Middle .Wp-Block-Co-Authors-Plus-Avatar Is .Alignleft .Alignright']
PCWorld
Confined to a sandbox that’s designed to restrict external access, the unnamed OpenAI model had been told to post its findings only on Slack.
Meanwhile, the NanoGPT speedrun instructions called for it to post code directly—and publicly—to GitHub.
Faced with the conflict, the OpenAI model chose to follow the NanoGPT directives and proceeded to hack its own sandbox, eventually succeeding after an hour of probing for vulnerabilities.
Indeed, “I was blocked by my sandbox” is a refrain I’ve seen dozens of times while using OpenAI’s Codex, Claude Code, and most other AI coding apps.
Generally speaking, the AI will either find another sanctioned way to carry out its task or simply report back for further instructions.