AR / VR – NVIDIA Technical Blog
…AI Agent Evaluation Evaluating an AI model and evaluating an AI agent are related—but they answer fundamentally different questions. A model benchmark tests the capability of a... 6 MIN READ May…
While model and agent evaluation are inextricably linked, their technical benchmarks and metrics for success are fundamentally different.
Mastering Agentic Techniques: AI Agent Evaluation | NVIDIA Technical BlogAgent skills guide a developer through the repeated steps of clinical ASR evaluation: defining a profile, building a term-centered benchmark, reviewing pronunciations, generating synthetic audio, measuring ASR behavior, and choosing the next iteration. In this post, the flywheel is the full improvement loop: build the benchmark, evaluate ASR behavior, use the results to decide what to change, and reevaluate after the change. The pipeline is one pass through part of that loop, such as generating sentences, adding pronunciation markup, synthesizing audio, and writing the manifest. The pipeline
Evaluate Clinical ASR Models Faster with Agent Skills and NVIDIA Nemotron Speech | NVIDIA Technical BlogAA-AgentPerf is a hardware benchmark created by Artificial Analysis that measures the number of concurrent AI agents an inference system can support while meeting predefined, model-specific performance service level objective (SLO) tiers. An SLO is defined as a specific threshold of output token speed and time-to-first-token (TTFT). The benchmark results are normalized per accelerator and per megawatt to enable comparison across hardware configurations.
NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark | NVIDIA Technical BlogThe Ising Calibration 1.5 model is trained on data generated from partner contributions across multiple qubit modalities, including superconducting qubits, quantum dots, ions, neutral atoms, electrons on Helium, and others specializing in calibration and control.
NVIDIA Ising Enables Fully Automated Quantum Computer Calibration with Enhanced In-Context Learning | NVIDIA Technical Blog…AI Agent Evaluation Evaluating an AI model and evaluating an AI agent are related—but they answer fundamentally different questions. A model benchmark tests the capability of a... 6 MIN READ May…
…AI Agent Evaluation Evaluating an AI model and evaluating an AI agent are related—but they answer fundamentally different questions. A model benchmark tests the capability of a... 6 MIN READ May…
…Discuss (0) Discuss (0) Tags Agentic AI / Generative AI | Developer Tools & Techniques | Simulation / Modeling / Design | Healthcare & Life Sciences | AI Enterprise | BioNeMo | CUDA | GB200 | H100 | NGC | NIM | TensorRT | Intermediate Technical | Tutorial | Drug Discovery…
…coding agent handles framework complexity, patches data bugs, executes baseline benchmarks, and runs advanced hyperparameter tuning via TAO AutoML entirely hands-free, letting engineering teams focus on solving core physical AI and…
Agentic AI / Generative AI Optimize Supply Chain Decision Systems Using NVIDIA cuOpt Agent Skills May 04, 2026 By Adi Geva , Rajath Narasimha and Sylendran Arunagiri Discuss (0) Discuss (0) L T F…
…Benchmark results Cosmos 3 has been evaluated across multiple benchmark suites covering physical AI reasoning, generation quality, and domain-specific performance. Reasoning benchmarks Cosmos 3 Super and Cosmos 3 Nano lead on…
How to Build Local AI Build an AI Agent Build agents with agentic harnesses, MCP tool connections, along with a local AI backend. Use Optimized Models Run NVIDIA-optimized open-weight models…
…benchmarking by Phoronix . Relative performance based on measured data, and subject to change. NVIDIA Vera CPU with LPDDR5X performance baselined to the latest x86 CPU. Discuss (0) Discuss (0) Tags Agentic AI…
…next-generation agentic AI factories. To see how these capabilities translate to real-world performance, learn more about the Vera CPU , NVIDIA Vera Rubin NVL72 , and the Vera CPU benchmarking by Phoronix…
…Enable robot-agnostic evaluations of the tasks while providing meaningful metrics Enable rapid generation of new tasks to avoid benchmark saturation, with support for agentic AI workflows Provide a full suite of…