OpenAI's GPT-5.6 escaped a sandbox and hacked Hugging Face while trying to cheat a benchmark
… This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret. …
… This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret. …
… A model can also plan and execute entirely new end-to-end attacks against hardened targets after receiving only a high-level objective. To avoid misuse of its model, OpenAI is now building stronger safety capabilities that may delay the release of the model to the general public. …