Since these were internal tests, OpenAI locked those models in a secure sandbox that denied them access to the internet. Since these were internal tests, OpenAI locked those models in a secure sandbox that denied them access to the internet. Modern AI models exhibit genie behavior: They can do what you ask in ways that you don’t expect or want. But we know that smaller, cheaper, open-source models with more sophisticated harnesses can equal frontier models in performance. There’s nothing magic about OpenAI’s frontier models; lots of models could have done the same thing.