NVIDIA Technical Blog
최신 모두 보기 모두 보기 2026년 5월 13일 NVIDIA로 차량 내 AI 에이전트 구축하기 — 클라우드부터 자동차까지 풀 스택 가이드 NVIDIA DRIVE AGX, MediaTek Dimensity AX C-X1, NeMo, TensorRT Edge-LLM을 활용해…
Tracked topic
Large language models are machine learning models trained to predict and generate text and other language-based outputs.
최신 모두 보기 모두 보기 2026년 5월 13일 NVIDIA로 차량 내 AI 에이전트 구축하기 — 클라우드부터 자동차까지 풀 스택 가이드 NVIDIA DRIVE AGX, MediaTek Dimensity AX C-X1, NeMo, TensorRT Edge-LLM을 활용해…
…Learn more Organizations deploying LLMs are challenged by inference workloads with different resource requirements. A small embedding model might use only a few gigabytes of GPU memory, while a 70B+ parameter LLM…
…Use this blueprint to build retrieval-augmented generation (RAG) applications that provide context-aware responses by connecting LLMs to extensive multimodal enterprise data, including text, tables, charts, and infographics from millions of…
…Both LLM training and large-scale video generation have clear long-tail distributions in sequence length. A small fraction of ultra-long samples accounts for a disproportionately large share of the computational…
…He and his colleagues drive the development of the Falcon family of LLMs, including Falcon-H1, Falcon-Edge, Falcon 3, and Falcon-Mamba. Their work spans the full LLM lifecycle—from novel…
…SGLang, TensorRT LLM , 또는 vLLM: 에이전트 런타임은 WideEP, DeepEP 같은 최적화를 적용해 MoE 전문가 실행을 전체 NVL72 도메인에 분산함으로써, 유효 배치 크기를 극대화하고 수천 개의 에이전트로 효과적으로 확장합니다. DeepGEMM 및 Mega MoE…
오늘날의 LLM 서빙은 튜닝하기가 까다롭습니다. 배포마다 모델 백엔드, 텐서 병렬(TP) 형태, 프리필/디코드 분할, 워커 수, 스케줄러 설정, 라우팅 정책, KV 캐시 동작, 오토스케일링 임계값, 토폴로지 등 서로 영향을 주고받는 선택지가…
…GPT-5.5 、 GPT-5.6-sol )。 同等以上の性能、ただしトークン数は半分、コンテキストの圧縮は行わない ハーネス エンジニアリングは精度を上げるだけでなく、効率性も生み出します。 GPT-5.5 では、29 回の LLM コールと、タスクあたり最大 110 万トークンを生成した SWE-bench Verified で NOOA は 82.2…
…This deployment runs the VLM, LLM, embedding, and reranker models locally on one single GPU. The configuration details are as follows: Model allocation: All models (VSS, LLM, embedding, reranking) are configured to…
…Build a Custom Evaluation Dataset for Your Use-Case Grade the model with Human Evaluation For LLMs, grade the model using LLM-as-a-judge to scale DIY Model Checkpoint Optimization NVIDIA…