LLMs and biorisk
…and mitigate the prospects for malicious actors to misuse the same capabilities that make AI so promising for scientists and innovators. When Anthropic released Claude Opus 4, we activated AI Safety Level…
…and mitigate the prospects for malicious actors to misuse the same capabilities that make AI so promising for scientists and innovators. When Anthropic released Claude Opus 4, we activated AI Safety Level…
…It is calculated based on the actor's activity across Claude.ai, Claude Code, and our API, drawing on our safety classifiers alongside open-source and internal threat-intelligence indicators. The higher…
…How our biology safeguards work One of the core ways we protect against misuse in biology is via safety classifiers: smaller, automated AI systems that detect when Fable 5 is asked to…
…and we therefore expect them to be motivated to try to circumvent our safety measures. Fable 5 comes with a new set of classifiers : separate AI systems that detect potential misuse, including…
…Sociotechnical alignment Which human values should AI models hold, and how should they operate in the face of conflicting or ambiguous values? How is AI used (and misused) in the wild? How…
…within 6 to 12 months, we expect that many other AI companies will have Mythos-class models, and they could release them without safeguards that prevent misuse. In that world, cyberattacks could…
…My secondary concern is the risk that powerful AI models may be misused to carry out cyberattacks or biological attacks, and may have serious alignment problems . Open-weights models—it does not…
Announcements Frontier threats red teaming for AI safety Jul 26, 2023 “Red teaming,” or adversarial testing, is a recognized technique to measure and increase the safety and security of systems. While previous…
…AI models, to begin to understand why the model behaves in certain ways—and, ultimately, to catch problems that standard testing might miss. A detailed description of all capability and safety evaluations…
…For many developers, the agentic AI era began with Sonnet-class models: Claude Sonnet 3.5, 3.6, and 3.7 were the first models that showed impressive skills in coding and…