Nvidia details Rubin architectural optimizations for inference – improvements target better performance and efficiency from the GPU to the rack
… Here are some key rates to keep in mind for this chip so far: Swipe to scroll horizontally Nvidia Rubin GPU Row 0 - Cell 1 NVFP4 Inference 50 PFLOPS with sparsity NVFP4 Training 35 PFLOPS FP8/FP6 Training 17.5 PFLOPS INT8 250 TOPS FP16/BF16 4 PFLOPS TF32 2 PFLOPS FP32 130 TFLOPS FP64 33 TFLOPS Let’… …