None
EN
OpenAI Tries to Train AI Not to Deceive Users, Realizes It’s Instead Teaching It How to Deceive Them While Covering Its Tracks
['Victor Tangermann', 'Senior Editor', 'Joe Wilkins', 'Frank Landymore', 'Noor Al-Sibai', 'Jon Christian']
Futurism
In reality, the team found, they were unintentionally teaching the AI how to more effectively deceive humans by covering its tracks.
As a result, AI models can end up deceiving the user, such as by claiming they’ve completed a task without ever having done so.
Researchers have long known that AI models are capable of deceiving the user to avoid perceived harm or pursue an obscure goal.
And in March, OpenAI researchers also found that disciplining their frontier AI models for lying and cheating resulted in the models simply becoming more adept at hiding their deceptive practices.
More on AI scheming: OpenAI Scientists’ Efforts to Make an AI Lie and Cheat Less Backfired Spectacularly
['tracks'
'scheming'
'users'
'wrote'
'efforts'
'covering'
'ai'
'models'
'openai'
'apollo'
'teaching'
'researchers'
'realizes'
'openais'
'model'
'train'
'deceive'
'instead'
'tries']