Search

Showing top 65 results for "model-by-model evaluation"

People also ask

Why build evaluations?

When teams first start building agents, they can get surprisingly far through a combination of manual testing, dogfooding, and intuition. More rigorous evaluation may even seem like overhead that slows down shipping. But after the early prototyping stages, once an agent is in production and has started scaling, building without evals starts to break down. The breaking point often comes when users report the agent feels worse after changes, and the team is “flying blind” with no way to verify except to guess and check. Absent evals, debugging is reactive: wait for complaints, reproduce manually

Demystifying evals for AI agents

What's next?

Claude Sonnet 4.5 represents a meaningful improvement, but we know that many of its capabilities are nascent and do not yet match those of security professionals and established processes. We will keep working to improve the defense-relevant capabilities of our models and enhance the threat intelligence and mitigations that safeguard our platforms. In fact, we have already been using results of our investigations and evaluations to continually refine our ability to catch misuse of our models for harmful cyber behavior. This includes using techniques like organization-level summarization to und

Building AI for cyber defenders

Followed topics

Search

People also ask

Claude Opus 4.6

Measuring LLMs’ ability to develop exploits

Introducing Claude Opus 4.7

AI agents find smart contract exploits

Automated Alignment Researchers: Using large language models to scale scalable oversight

Natural Language Autoencoders

Introducing Claude Opus 4.5

Measuring LLMs' impact on N-day exploits

Building Effective AI Agents

Claude Fable 5 and Claude Mythos 5