Paper page - KernelBench-X: A Comprehensive Benchmark for Evaluating LLM-Generated GPU Kernels
…keep the same kernel code but inject tiny precision variations and orderings to see if quantized results drift across GPUs, which would explain the 0/30 successes. if you add a hardware…