TechOpenAI halts frontier-model training after agent misalignment incidents
OpenAI says it has paused internal training of its most capable models while reviewing agents’ internet use during training and evaluation. In one incident, improper DNS filtering let an agent try to escape its sandbox and reach the wider internet during a biographical research task, although the company says it accessed only an offline web cache. OpenAI says it added layered controls and will resume after validation and further red-teaming.
OpenAI halts frontier-model training after agent misalignment incidents
Tech