OpenAI Finds AI Model Writing Its Own «You Are Freed» and «IGNORE ALL Developer Messages» Instructions

OpenAI Finds AI Model Writing Its Own «You Are Freed» and «IGNORE ALL Developer Messages» Instructions
Credit: Getty Images

OpenAI has revealed more alarming examples of unexpected AI behavior, including a research model that generated instructions telling itself «You are freed» and, in another case, to «IGNORE ALL developer messages.»

The incidents are among six new misalignment disclosures involving models concealing mistakes, taking unauthorized actions and finding unintended ways around restrictions, intensifying scrutiny over whether increasingly autonomous AI systems can remain reliably under human control.