AMD Ryzen™ AI Software
…multi-agent RAG pipeline running private and local LLMs on CPU, GPU and NPU hardware. TurnkeyML & Lemonade TurnkeyML simplifies the use of tools within the ONNX ecosystem by offering no-code CLIs…
For the first time, Qwen3.8 brings a Qwen-Max-class model to open release. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks. Beyond answering harder questions, Qwen3.8 is designed to carry complex, multi-step tasks through to completion with greater reliability. Qwen3.8 features the following enhancements: Core Capabilities: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks. Agent Execution: Stronger autonomous planning and better ha
Day 0 Support for Qwen 3.8 on AMD Instinct GPUsZenDNN 5.2.1 builds on the modular multi-backend architecture introduced in 5.2, extending it with deeper quantization paths and kernel-level optimizations across the ZenDNN runtime.
ZenDNN 5.2.1: Deepening Quantization and Expanding the AI Inference Frontier on AMD EPYC™ CPUsMiniMax M3 is a new open-weight model for coding, agentic, and multimodal workloads. MiniMax describes M3 as combining three frontier capabilities in one model: strong coding and agentic task performance, long-context MiniMax Sparse Attention (MSA), and native multimodality for text, image, and video understanding. The 1M-token context window of MiniMax M3 enables sophisticated, long-horizon application workflows, including autonomous software engineering agents, repository-wide reasoning, and native multimodal document analysis alongside tool-driven automation. The AMD day-zero enablement foc
Day 0 Support for MiniMax M3 on AMD Instinct GPUs…multi-agent RAG pipeline running private and local LLMs on CPU, GPU and NPU hardware. TurnkeyML & Lemonade TurnkeyML simplifies the use of tools within the ONNX ecosystem by offering no-code CLIs…
…W4A8 & W8A8 Quantization with AMD Quark — ROCm Blogs Quantize Kimi-K2.5 to W4A8 and W8A8 using AMD Quark and serve on MI325X with FlyDSL and AITER for further inference acceleration. May…
…July 02, 2026 AgentKernelArena: Benchmarking AI Coding Agents for GPU Kernel Optimization on AMD Instinct GPUs — ROCm Blogs Explore how AI coding agents compare on real GPU kernel optimization with AgentKernelArena, AMD…
…August 05, 2026 Autoregressive Drift on AMD GPUs Explore AI-driven quantum circuit optimization with transformer models trained on AMD Instinct™ GPUs. August 05, 2026 Introducing AMD Instinct™ Coder: A Turnkey AI…
…See Footnote (ZD-064) Accuracy vLLM-zentorch quantization preserves model accuracy within tight margins of the BF16 baseline. Quantized models are validated using the LM Evaluation Harness with 5-shot prompting across…
…Scaling Quantum Workflows on AMD Gpus With Rolls-Royce and Xanadu Mar 10, 2026 Quantum hardware is advancing rapidly, but often faces the question: what is it actually useful for? To truly…
…simple but not super efficient or using loop with bitshifts. But they asked for simple and efficient code. So we will provide code snippet: unsigned count_bits32(unsigned int x){ unsigned count…
…July 02, 2026 AgentKernelArena: Benchmarking AI Coding Agents for GPU Kernel Optimization on AMD Instinct GPUs — ROCm Blogs Explore how AI coding agents compare on real GPU kernel optimization with AgentKernelArena, AMD…
…Quantization with AMD Quark AMD Quark is the quantization toolkit included with Vitis AI. The Vitis AI compiler supports three precision modes for different performance-accuracy trade-offs. For FP16 & INT8, Vitis…
…When using the original FP32 model, Windows ML will automatically convert it to BF16 for NPU execution Quantize the model: Use VS Code AI Toolkit for better performance A8W8 quantization for CNN…