Improving Fable 5 Safeguards
… As new biological capabilities develop on the frontier of AI, we need to be cautious to ensure that the new risks they pose do not materialize ahead of their potential scientific benefits. …
… As new biological capabilities develop on the frontier of AI, we need to be cautious to ensure that the new risks they pose do not materialize ahead of their potential scientific benefits. …
… In new research carried out with collaborators at AE Studio, we explore a new method that could enable the benefits of training many separately-filtered models, but at the cost of training only one model. …
… For the model’s agentic and computer use capabilities, we’ve also made considerable progress on defending against prompt injection attacks, one of the most serious risks for users of these capabilities. …
… We’ll be tracking these changes over time as part of our Economic Index as model capabilities, products, and adoption continue to progress. These productivity gains come from making existing tasks faster to complete. …
… We tested both production models and “helpful-only” variants that were trained differently. We also tested some base pretrained models before post-training. We found that post-training significantly impacts introspective capabilities. …
… We bring the industry context, the engineering scale and the trust frameworks that use Claude to deliver production outcomes inside the most demanding enterprise environments. …
… Finally, Claude Code’s default settings push new users towards approval-based oversight since actions are not auto-approved by default , so some of the shifts we observe may reflect Claude Code’s product design. …
… The people closest to product requirements and users are best positioned to define success. With current model capabilities, product managers, customer success managers, or salespeople can use Claude Code to contribute an eval task as a PR—let them! …
… We’re working as quickly as we can to safely release Mythos-level capabilities in general access. To do so, we’ll need highly robust safeguards that prevent the model’s cyber capabilities from being misused—safeguards that we and, to our knowledge, all other AI developers have yet to develop. …
… Each model brings a new set of capabilities that are sometimes put to work in unexpected ways. …