Introducing Claude Sonnet 5
… More recently, though, the clearest gains in agentic capabilities have been in our Opus-class models. Sonnet 5 narrows the gap: its performance is close to that of Opus 4.8, but at lower prices. …
… More recently, though, the clearest gains in agentic capabilities have been in our Opus-class models. Sonnet 5 narrows the gap: its performance is close to that of Opus 4.8, but at lower prices. …
… We cover the model, our new product updates, our evaluations, and our extensive safety testing in depth below. First impressions We build Claude with Claude. Our engineers write code with Claude Code every day, and every new model first gets tested on our own work. …
… OSWorld-Verified : We made changes to how we run the OSWorld-Verified evaluation in order to more accurately reflect the model’s performance in the real world, and have updated the Opus 4.7 score to 82.3%. Read more about the updates in the System Card . …
… Each of these updates takes advantage of Claude Opus 4.5’s market-leading performance in using computers, spreadsheets, and handling long-running tasks. …
… As noted above, it hasn’t been applied to any of our Claude models. Our evaluations quantify performance in terms of next-token prediction ability, rather than performance on real downstream tasks. …
… This tallies with external testers’ experience of Mythos Preview’s performance, and with recent additional evaluations of the model: The UK’s AI Security Institute reports that Mythos Preview is the first model to solve both of their cyber ranges simulations of multistep cyberattacks end to end; Mo… …
… Performance that would have previously required reaching for an Opus-class model—including on real-world, economically valuable office tasks —is now available with Sonnet 4.6. The model also shows a major improvement in computer use skills compared to prior Sonnet models. …
Interpretability A “diff” tool for AI: Finding behavioral differences in new models Mar 13, 2026 Read the paper Every time a new AI model is released, its developers run a suite of evaluations to measure its performance and safety. …
… It’s the new default model on Claude Max, and the strongest model on Claude Pro. Performance and cost-effectiveness Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. …
… When we had the model play the deck-building game Slay the Spire , giving it access to persistent file-based memory improved its performance three times more than for Opus 4.8; Fable also reached the game’s final act three times more often. …