What is AI Networking?
…In AI clusters, these capabilities are complementary. Article By AMD Networking Related Blogs View All Blogs QuickReduce FP4 Quantization and Benchmarking on MI355 — ROCm Blogs Learn how QuickReduce uses FP4 quantization to…
…In AI clusters, these capabilities are complementary. Article By AMD Networking Related Blogs View All Blogs QuickReduce FP4 Quantization and Benchmarking on MI355 — ROCm Blogs Learn how QuickReduce uses FP4 quantization to…
…May 21, 2026 QuickReduce FP4 Quantization and Benchmarking on MI355 — ROCm Blogs Learn how QuickReduce uses FP4 quantization to accelerate all-reduce communication and evaluate its performance on AMD Instinct MI355 GPUs…
…W4A8 & W8A8 Quantization with AMD Quark — ROCm Blogs Quantize Kimi-K2.5 to W4A8 and W8A8 using AMD Quark and serve on MI325X with FlyDSL and AITER for further inference acceleration. May…
…May 26, 2026 QuickReduce FP4 Quantization and Benchmarking on MI355 — ROCm Blogs Learn how QuickReduce uses FP4 quantization to accelerate all-reduce communication and evaluate its performance on AMD Instinct MI355 GPUs…
…W4A8 & W8A8 Quantization with AMD Quark — ROCm Blogs Quantize Kimi-K2.5 to W4A8 and W8A8 using AMD Quark and serve on MI325X with FlyDSL and AITER for further inference acceleration. May…
…But benchmarks don't answer the question that actually matters: Should you undertake this effort and is it viable for your business? In this interactive technical discussion, we’ll break down the…
…This session covers deployment strategies, ZenDNN acceleration, Venice quantization support, kernel optimizations, and EPYC’s role in powering Agentic AI workloads. July 22, 2026 15:30 - 15:50 Senior Director, Software Engineering…
…Benchmarking AI Coding Agents for GPU Kernel Optimization on AMD Instinct GPUs — ROCm Blogs Explore how AI coding agents compare on real GPU kernel optimization with AgentKernelArena, AMD's open benchmarking arena…
…Local Benchmarks Definition of segment in these tests. In the context of the local benchmarks, a segment is one bounded execution unit of TGP: the model decodes a limited portion of the…
…Quantization Support : Experimental INT4 support for LLMs and specialized UINT4/W8A8 quantization for recommendation systems (DLRM-v2). BFloat16 & Graph Optimizations : Enhanced EPYC™ processor specific kernels and advanced pattern identification ensure that every…