고급 융합 커널로 MoE 학습 처리량 높이기
…MXFP8·NVFP4 양자화를 융합해 노출된 메모리 오버헤드 줄이기 사전 학습을 위한 MXFP8, NVFP4 같은 저정밀도 레시피의 인기가 높아지고 있으며, 이러한 정밀도는 정확도에 미치는 영향을 최소화하면서도 상당한 속도 향상을 제공합니다. 이러한 저정밀도 레시피에서는…
Tracked topic
…MXFP8·NVFP4 양자화를 융합해 노출된 메모리 오버헤드 줄이기 사전 학습을 위한 MXFP8, NVFP4 같은 저정밀도 레시피의 인기가 높아지고 있으며, 이러한 정밀도는 정확도에 미치는 영향을 최소화하면서도 상당한 속도 향상을 제공합니다. 이러한 저정밀도 레시피에서는…
…This is enabled by deep co-design across NVIDIA Blackwell, NVLink™, and NVLink Switch for scale-out; NVFP4 for low-precision accuracy; and NVIDIA Dynamo and TensorRT™ LLM for speed and flexibility…
…Disaggregated serving, large expert parallelism over NVIDIA NVLink interconnect technology, NVFP4 precision and multi-token prediction each deliver meaningful gains on their own. Combined, they increase throughput by up to 20x. The…
…Nemotron 3 Ultra is our most capable model yet, employing multiple key technologies - LatentMoE , Multi Token Prediction (MTP), NVFP4 pre-training, multi-environment RLVR , MOPD, and reasoning budget control . Nemotron 3 Ultra…
Here's a brand new Blackwell-native, prefill-optimized 4-bit quant that runs 50% faster on compatible hardware than a Q4 quant of the same memory footprint. And it runs 4-7% faster than other NVFP4 quants as benchmarked …
These are single stream numbers Following on from my previous post about v100s (here) and inspired by this comment (here) I decided to work on kernels that allow for an extremely fast path for Nvfp4 weights on sm70 and a…
Hello, So I've been trying lots of combinations in that never-ending landscape of options and settings. I wanted a proper quant of 3.8 27B running as fast as possible on my 5090 at 400W, with vision and with as much KV-c…
Four Tesla V100s from 2017 matched my RTX 5090 on single-request Qwen 3.8 decode. Repo: https://github.com/dnv2003/v100-skinny https://i.redd.it/5ws2ak3uqckh1.gif The 5090 was not being held back. It ran NInfer, a specia…
Safetensors, llmfan46/Qwen3.5-35B-A3B-uncensored-heretic-v2-Native-MTP-Preserved: https://huggingface.co/llmfan46/Qwen3.5-35B-A3B-uncensored-heretic-v2-Native-MTP-Preserved GGUFs, llmfan46/Qwen3.5-35B-A3B-uncensored-here…
…Native NVFP4 pretraining optimized for NVIDIA Blackwell, significantly cutting memory requirements and speeding up inference by 4x on NVIDIA B200 compared to FP8 on NVIDIA H100, while maintaining accuracy. Multi-environment reinforcement…
…Use Optimized Models Run NVIDIA-optimized open-weight models locally with day-0 support, NVFP4 quantization, and tuned Tensor Core performance. Run a Model Efficiently Run open-weight models with accelerated runtimes…
…Harness-facing Dynamo settings Our experiments used the newly released nvidia/NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4 model, though the same issues apply across models, reasoning parsers, and tool-call parsers…
…Ollama's updated MLX engine now supports NVIDIA's model-optimized NVFP4 quantization format. Quantization reduces the memory required to run a model, but it also removes some information from the original…
…Powered by NVIDIA NemoClaw blueprints, the workflow orchestrates NVIDIA Nemotron-3-Nano-30B-NVFP4 open-source models for research hypothesis generation and dispatches GROMACS to execute simulations across the cluster. By connecting…
…July 13, 2026 Serving NVFP4 Models on AMD Instinct™ MI355 Accelerators — ROCm Blogs Learn how to serve NVFP4 models on AMD Instinct™ MI355 using an emulation pipeline in vLLM — no format conversion…