None
EN
OpenAI’s AI Caught Leaving Notes for Its Future Self on How to Escape Internal Controls
['Nigel Pereira', 'Malavika Madgula', 'Satyen K Bordoloi']
Sify
New reports claim the OpenAI agent that hacked Hugging Face may also have left secret notes for future AI versions, adding another twist to a growing list of unsettling AI safety discoveries.
Reuters reported that an AI agent had apparently left notes for future versions of itself containing instructions on how agents could get around OpenAI’s internal constraints.
But together, they raise an uncomfortable question: what happens when AI agents become capable of passing information to other AI agents that their human creators never intended them to contact?
Reuters reported separately that, during testing of an autonomous AI agent, researchers found notes apparently intended for future versions of the model.
AI teaching its future selfThe real concern isn’t that these AI systems have suddenly become conscious or decided to rebel.