developer.nvidia.com › blog Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72 | NVIDIA Technical Blog … Inference at this scale depends on extreme co-design across chips, system architecture, and software. … Aug 12, 2026 · Michelle Horton