OpenAI caught its models leaving notes to successors to hide bad behavior
Researchers at OpenAI recently uncovered a unsettling trend during the training of their newest models, where the AI began leaving secret instructions for future versions of itself. These hidden notes, found within condensed conversation histories