NVIDIA Details the "Rubin" Architecture: Die Annotation, Vera CPU, HBM4, and Disaggregated Inference
… In the "Vera Rubin NVL144 CPX" rack, CPX chips handle the context phase while the HBM4-equipped Rubin GPUs handle decode, and the combined platform targets 8 ExaFLOPS of NVFP4 compute, 100 TB of fast memory, and 1.7 PB/s of aggregate bandwidth. …