A “diff” tool for AI: Finding behavioral differences in new models
… Acknowledgements This post was authored by Thomas Jiralerspong Anthropic Fellows Program and Trenton Bricken Anthropic Alignment Science . …
Tracked topic
Stay informed on Anthropic’s AI development, focusing on Claude AI updates, safety protocols, and the future of AGI threats.
… Acknowledgements This post was authored by Thomas Jiralerspong Anthropic Fellows Program and Trenton Bricken Anthropic Alignment Science . …
… It was developed as a collaboration between UC Berkeley, the Max Planck Institute for Security and Privacy, UC Santa Barbara, and Arizona State University with contributions from security researchers at Anthropic, OpenAI, and Google , as a follow-on to the CyberGym vulnerability-reproduction benchm… …
… Related content Apply for Anthropic’s AI for Science rare disease research grants Anthropic is sharing a focused call for AI for Science applications centered specifically on rare genetic diseases. …
… Model Provider Adapter Claude Opus 4.7 Anthropic claude agent sdk Claude Opus 4.6 Anthropic claude agent sdk Claude Mythos Preview Anthropic claude agent sdk Claude Opus 4.5 Anthropic claude agent sdk Claude Opus 4.1 / 4 Anthropic claude agent sdk GPT-5.4 OpenAI via OpenRouter openrouter GPT-5.1 Op… …
… Claude Opus 4.6 is the best Anthropic model we’ve tested. …
… Donating our open-source alignment tool Focus areas for The Anthropic Institute At The Anthropic Institute TAI , we’ll be using the information we can access from within a frontier lab to investigate AI’s impact on the world, and sharing our learnings with the public. …
… All runs used the Terminus-2 harness, except for OpenAI’s Codex CLI. …