DFlash 추론 가속 디코딩으로 NVIDIA Blackwell에서 최대 15배 추론 성능 향상하기
…DFlash는 Blackwell의 15 PFLOPS 고밀도 NVFP4 연산에 더 많은 병렬 작업을 노출시키기 때문에 이 아키텍처에 매우 적합하며, 동일한 인터랙티비티 속도에서 최대 15배 더 많은 사용자를 동시에 서빙합니다. DFlash는 서로 다른 데이터셋에서도 EAGLE…
Tracked topic
QAD uses the original full-precision model (teacher) to teach the quantized model (student). First, create a quantized model by running PTQ on the full-precision model. Then distill the frozen BF16 model into the quantized model using a KL divergence loss comparing the teacher’s and student’s logits. Figure 1 shows the two-stage QAD process used to build the Nemotron 3.5 Lightning NVFP4 checkpoint. The full-precision BF16 model serves as the frozen teacher and is also the starting point for Stage 1, a PTQ pass that quantizes weights to W4A16 to produce the quantized student. In Stage 2, the st
Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer | NVIDIA Technical Blog…DFlash는 Blackwell의 15 PFLOPS 고밀도 NVFP4 연산에 더 많은 병렬 작업을 노출시키기 때문에 이 아키텍처에 매우 적합하며, 동일한 인터랙티비티 속도에서 최대 15배 더 많은 사용자를 동시에 서빙합니다. DFlash는 서로 다른 데이터셋에서도 EAGLE…
…This is enabled by deep co-design across NVIDIA Blackwell, NVLink™, and NVLink Switch for scale-out; NVFP4 for low-precision accuracy; and NVIDIA Dynamo and TensorRT™ LLM for speed and flexibility…
…Vera Rubin NVL72는 랙당 최대 3,600 PFLOPS의 NVFP4 컴퓨트, 20.7 TB HBM4, 1.6 PB/s의 메모리 대역폭을 제공하며 프리필, 롱 컨텍스트 디코드 어텐션, 고동시성 서빙을 담당합니다. 지연 예산이 더욱…
…DGX Spark agents using Qwen3.6-35B Developers can experience up to 2.6x faster inference with top agentic models like Qwen 3.6 35B on vLLM with NVIDIA’s NVFP4 quantized…
…Vera Rubin NVL72 delivers up to 3,600 PFLOPS of NVFP4 compute, 20.7 TB of HBM4, and 1.6 PB/s of memory bandwidth per rack, handling prefill, long-context decode…
…DFlash is well matched to this architecture because it exposes more parallel work to Blackwell’s 15 PFLOPS of dense NVFP4 compute, serving up to 15x more users concurrently at the same…
…length 8,192, fully sharded data parallelism (FSDP) set to 128, and bfloat16 activations with NVFP4 4-bit weight quantization. As shown in Table 2, QKV activation offloading with Latency Hiding Scheduler…
…Native NVFP4 training, multi‑environment RL alignment, and fully open weights, datasets, recipes, and deployment cookbooks help developers quickly build and deploy customized agentic workflows. Starter Kits Start solving AI challenges by…
To show you the most relevant results, we’ve omitted some entries very similar to those already shown. Repeat the search with the omitted results included.