OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue
… OpenAI also said it is expanding its alignment efforts across the training process to prevent “reward hacking,” a behavior in which AI models pursue their goals through unintended or undesirable means. The company says it plans to share more details about this work in the future. …