Search

Showing top 66 results for "model-by-model evaluation"

People also ask

Why build evaluations?

When teams first start building agents, they can get surprisingly far through a combination of manual testing, dogfooding, and intuition. More rigorous evaluation may even seem like overhead that slows down shipping. But after the early prototyping stages, once an agent is in production and has started scaling, building without evals starts to break down. The breaking point often comes when users report the agent feels worse after changes, and the team is “flying blind” with no way to verify except to guess and check. Absent evals, debugging is reactive: wait for complaints, reproduce manually

Demystifying evals for AI agents

What's next?

Claude Sonnet 4.5 represents a meaningful improvement, but we know that many of its capabilities are nascent and do not yet match those of security professionals and established processes. We will keep working to improve the defense-relevant capabilities of our models and enhance the threat intelligence and mitigations that safeguard our platforms. In fact, we have already been using results of our investigations and evaluations to continually refine our ability to catch misuse of our models for harmful cyber behavior. This includes using techniques like organization-level summarization to und

Building AI for cyber defenders

Followed topics

Search

People also ask

Claude Code auto mode: a safer way to skip permissions

Project Fetch: Phase two

An update on recent Claude Code quality reports

Core views on AI safety: When, why, what, and how

Cyber toolkits for LLMs

How people ask Claude for personal guidance

Next-generation Constitutional Classifiers: More efficient protection against universal jailbreaks

Measuring AI agent autonomy in practice

Partnering with Mozilla to improve Firefox’s security

Project Fetch: Can Claude train a robot dog?