AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

RLVR is a Kernel, Not a Function: Statistical Inference for pass@$k$ Crossovers

arXiv · AI, language, vision and robotics · article · Sep 18, 2026 · UTC

Reinforcement learning with verifiable rewards (RLVR) often improves pass@1 while falling behind its base model at larger sampling budgets $k$, a crossover read as evidence that RLVR only sharpens existing capability. We identify two limits to this reading. First, a visible crossing need not be statistically established: comparing models on the same prompts, we build confidence bands across sampling budgets $k$ that require evidence of both an early gain and a later loss. Across five public RLVR pairs no crossing is statistically established in the initial evaluations, while a 32k-token evalua

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T12:01:45.602Z. This is not the publication date.