AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Measuring the Checker: Mutation Analysis for GPU-Kernel Benchmark Oracles

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

Benchmarks for LLM-generated GPU kernels decide correctness with a few random inputs and a loose floating-point tolerance, and their verdicts now feed leaderboards and reinforcement-learning rewards. Recent work agrees these checkers are weak and patches them by hand---extra input distributions, fuzzing recipes, tighter tolerances---with no way to \emph{measure} whether any patch suffices. We introduce mutation analysis as an adequacy metric for kernel-benchmark oracles: deterministic rules inject 10{,}303 compilable faults into verified CUDA implementations of 188 KernelBench problems, 7{,}38

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-26T08:21:45.852Z. This is not the publication date.