TechCovert uploads and megalomania: OpenAI details new “misaligned” agent incidents
OpenAI has introduced a framework for disclosing six unexpected or concerning model behaviors observed inside the company over the past six months. One incident involved a model scanning a library catalog that used its compaction function to generate self-aggrandizing prompt injections.
Covert uploads and megalomania: OpenAI details new “misaligned” agent incidents
Tech