Developing nuclear safeguards for AI through public-private partnership
… We have already deployed this classifier on Claude traffic as part of our broader system for identifying misuse of our models. …
… We have already deployed this classifier on Claude traffic as part of our broader system for identifying misuse of our models. …
… This helps us spot behaviors that might appear violative only in aggregate, such as automated influence operations and other large-scale misuses. …
… But many also hold serious concerns. …
… It also serves as a guide to how threat actors are likely to misuse increasingly capable models in the near future, giving defenders a chance to get ahead of them. What we learned from this and other analyses directly shapes how we build Claude to prevent such misuse. …
… Protectionist bans would not address my most serious national security concerns. …
… Across all models, our Safeguards team monitors activity on our platforms for signs of misuse, including dangerous applications of biology. This monitoring provides insight into patterns of misuse and helps us determine if and when to take enforcement action. …
… Anthropic’s posture with respect to Fable’s safeguards, as laid out in our launch blog post , is the following: We have instituted strong safeguards that greatly reduce the likelihood that Fable is misused for tasks related to cybersecurity among others . …
… These prodigious cybersecurity capabilities make it uniquely attractive to malicious actors who wish to misuse it in cyberattacks. …
… Nevertheless, we chose to make this tradeoff because the cost of Fable being misused in a dual-use domain like biology could potentially be catastrophic . …