DeepStream SDK
…the performance using DeepStream, check the documentation . Read Customer Stories Enhance City Safety and Mobility The City of Raleigh used NVIDIA DeepStream and the VSS Blueprint to build AI agents and provide…
With agents running 24 hours a day, seven days a week on increasingly complex tasks, efficient local compute matters even more. NVIDIA has collaborated with the open source community to enhance the top inference backends for agents, llama.cpp and vLLM. llama.cpp now delivers 2x performance on Qwen 3.5 and 3.6 27B dense models, and 1.6x performance on Qwen 3.5 and 3.6 35B mixture-of-expert (MoE) models. The following two techniques make this possible: Multi-Token Prediction (MTP): An advanced speculative decoding technique, where a smaller draft model proposes several tokens ahead that the targ
Build Personal AI Agents on Windows PCs with New Tools from Microsoft and NVIDIA | NVIDIA Technical BlogAA-AgentPerf is a hardware benchmark created by Artificial Analysis that measures the number of concurrent AI agents an inference system can support while meeting predefined, model-specific performance service level objective (SLO) tiers. An SLO is defined as a specific threshold of output token speed and time-to-first-token (TTFT). The benchmark results are normalized per accelerator and per megawatt to enable comparison across hardware configurations.
NVIDIA Achieves Leading Agentic Coding Performance on First Agentic AI Benchmark | NVIDIA Technical BlogEarlier this week at GTC Taipei, NVIDIA unveiled the NVIDIA RTX Spark product family, including small form factor desktops and laptops built for the age of personal assistants. These desktops and laptops deliver 1 petaflop of AI power, up to 128 GB of memory, and CUDA-accelerated AI frameworks for running large models alongside everyday work. Microsoft is creating an RTX Spark special developer edition—the Microsoft Surface NVIDIA RTX Spark Dev Box—preloaded with a modified Windows configured for developers and the top developer tools you need to get started. To learn more, see Building the n
Build Personal AI Agents on Windows PCs with New Tools from Microsoft and NVIDIA | NVIDIA Technical BlogOne popular way to run AI locally has been to use multiple GPUs to access more memory and compute. While cloud frameworks like vLLM are well optimized for multiple GPUs thanks to their use in data centers, PC frameworks like llama.cpp and the ComfyUI implementation in PyTorch are not optimized for it. To solve this challenge, NVIDIA has collaborated with both llama.cpp and ComfyUI to enhance performance for RTX PCs with two equivalent GPUs. This enables you to run larger models and use the compute of both GPUs for better performance. llama.cpp now supports tensor parallelism (TP), fully utiliz
Build Personal AI Agents on Windows PCs with New Tools from Microsoft and NVIDIA | NVIDIA Technical Blog…the performance using DeepStream, check the documentation . Read Customer Stories Enhance City Safety and Mobility The City of Raleigh used NVIDIA DeepStream and the VSS Blueprint to build AI agents and provide…
…NVIDIA NeMo NVIDIA NeMo™ is an open suite of libraries for building, customizing, evaluating, and governing AI agents, helping developers optimize specialized models and agentic workflows across cloud, data center, and hybrid…
…The combination of Vera Rubin NVL72 and LPX enables this heterogeneous architecture, pairing large-scale AI factory performance with the fast token generation needed to power continuously running agentic systems and next…
…with full ConnectX-7 performance. Start building on NVIDIA DGX Spark → Discuss (0) Discuss (0) Tags Agentic AI / Generative AI | General | DGX | Intermediate Technical | Deep dive | AI Agent | Computex 2026 | DGX Spark…
…AI agent lifecycle management by fine-tuning, deploying, and continuously optimizing Nemotron models with NVIDIA NeMo™. NVIDIA TensorRT-LLM TensorRT™-LLM is an open-source library built to deliver high-performance, real…
…Kao Discuss (0) Discuss (0) L T F R E AI-Generated Summary Like Dislike Customizing AI agents enhances their performance on specialized tasks by refining reasoning, tool selection, output structure, and…
…define the question, set the budget, decide which mutations are allowed, and review the ledger, while the AI agent performs the repetitive work of trying bounded candidate strategies and recording the results…
…developer productivity and the performance and portability of HPC applications. NVIDIA NeMo Build, Customize, and Deploy Generative AI Models Segments : Banking, Payments, Trading NVIDIA NeMo™ is an agent-first, open suite of…
Agentic AI / Generative AI NVIDIA Nemotron 3 Nano Omni Powers Multimodal Agent Reasoning in a Single Efficient Open Model Simplify pipelines and improve multimodal reasoning accuracy with NVIDIA Nemotron 3 Nano Omni…
…AI-Q can read enterprise data, perform retrieval and synthesis, and create reports without raw documents leaving the controlled environment. This is critical for enterprises with data sovereignty requirements. The agent harness…