Trustworthy agents in practice
…Open protocols also keep competition focused on the quality and safety of the agent, rather than on who controls the integrations. None of these measures replace the work that model developers have…
…Open protocols also keep competition focused on the quality and safety of the agent, rather than on who controls the integrations. None of these measures replace the work that model developers have…
…Such a framework would allow AI developers to speak to governments (and vice versa) in consistent terms about the risks posed by each jailbreak. What we’re sharing today reflects our current…
…We’re also introducing adaptive thinking , where the model can pick up on contextual clues about how much to use its extended thinking, and new effort controls to give developers more control…
…Because it’s trained on a fixed dataset, Claude has a “knowledge cutoff,” so it won’t automatically know about recent developments like candidate announcements, media coverage, or election results. But when…
…Improvements in consistency, instruction following, and more have made developers with early access prefer Sonnet 4.6 to its predecessor by a wide margin. They often even prefer it to our smartest…
…AI-assisted capability development The most common technique family we observed was ATT&CK ID T1587 (Develop Capabilities), used by 574 of the 832 actors in our analysis, or 69%. The majority…
…PwC is reinventing how it executes deals end-to-end — diligence, value creation, integration — with agents working alongside deal teams. For private equity sponsors and corporate acquirers, this compresses the path from…
…use on the Claude Developer Platform The future of AI agents is one where models work seamlessly across hundreds or thousands of tools. An IDE assistant that integrates git operations, file manipulation…
…In early testing, we’re seeing the potential for a significant leap for our developers with Claude Opus 4.7. It catches its own logical faults during the planning phase and accelerates…
…To check this, we tested Claude's time estimation capabilities against a dataset of thousands of real-world software development tasks gathered from JIRA tickets for open-source repositories, with both developer…