None
EN
Guardrails For AI Didn’t Work. So Bring In The Philosophers
['Bruce Barcott']
ArtsJournal
When a powerful new AI model escapes the bounds of its developer’s sandbox —the entire purpose of which is to keep it in—and immediately goes Viking berserker, it’s bad.
Funny enough, people inside and outside of the tech world have been warning us about this for years.
Even funnier: A few days later Anthropic’s new AI model did the same thing, hacking into three organizations during a similar test run.
In the AI world the word they’ve developed for this is alignment.
So when OpenAI’s latest model goes rogue and starts tunneling into the digital version of the First National Bank, it should surprise no one.