NVIDIA Brings Zero-Code-Change Acceleration, IO Performance Gains, and Out-of-Core XGBoost | NVIDIA Technical Blog
… Please note that in this case, the “Kernel decompress” condition uses a CUDA async pool MR, the “Hardware decompress” uses an RMM pool MR, and the performance improvement in “Hardware decompress” is due to the low latency and high throughput of the Blackwell Decompression Engine. …