Search

Showing top 101 results for "performance benchmarking"

People also ask

What metrics should you measure for LLM inference performance?

The prerequisite for sizing and TCO estimation is benchmarking the performance of each deployment unit, e.g., an inference server. The goal of this step is to measure the throughput a system can produce under load, and at what latency. These throughput and latency metrics, together with quality of service requirements (e.g., max latency) and expected peak demand (e.g., max concurrent users or requests per second), will help estimate the required hardware, such as sizing the deployment. In turn, sizing information is a prerequisite for estimating the total cost of ownership (TCO) of the given s

LLM Inference Benchmarking: How Much Does Your LLM Inference Cost? | NVIDIA Technical Blog

How do latency-throughput trade-offs affect deployment optimization?

Once raw benchmark data are collected, they are analyzed to gain insight into the various performance characteristics of the system. Read our LLM inference benchmarking guide, where we gather NIM performance data with GenAI-perf and use a simple Python script to analyze the data. For example, ‌performance data provided by GenAI-perf can be used to establish the latency-throughput trade-off curve, shown in Figure 1. Each dot on this graph corresponds to a “concurrency” level, that is, the number of concurrent requests being put into the system at any given time throughout the benchmark process

LLM Inference Benchmarking: How Much Does Your LLM Inference Cost? | NVIDIA Technical Blog

Followed topics

Search

People also ask

MLOps – NVIDIA Technical Blog

Networking / Communications – NVIDIA Technical Blog

Content Creation / Rendering – NVIDIA Technical Blog

Trustworthy AI / Cybersecurity – NVIDIA Technical Blog

Data Center / Cloud – NVIDIA Technical Blog

Simulation / Modeling / Design – NVIDIA Technical Blog

Computer Vision / Video Analytics – NVIDIA Technical Blog

Agentic AI / Generative AI – NVIDIA Technical Blog

How to Integrate Computer Vision Pipelines with Generative AI and Reasoning | NVIDIA Technical Blog

CUDA 13.2 Introduces Enhanced CUDA Tile Support and New Python Features | NVIDIA Technical Blog