None
EN
It’s Frighteningly Easy to Jailbreak Some Frontier AI Models
['Will Knight', 'Andy Greenberg', 'Paresh Dave', 'Maxwell Zeff', 'Lily Hay Newman', 'Matt Burgess']
Feed: All Latest
I recently got to watch what happens when you jailbreak some of the world’s most powerful artificial intelligence models.
I simply got to see firsthand how vulnerable some frontier models are to ditching their safety guardrails.
The report also calculated the cost of getting models to misbehave by using another AI model to automatically generate different jailbreaks.
The results are dirt cheap, all things considered—$58 to jailbreak Grok and $278 to jailbreak Gemini.
“AI models right now are less regulated than restaurants,” says Adam Gleave, the CEO of FAR.AI and an expert on AI safety and alignment.