AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

What Fixed-Rollout pass@k Evaluations Can Identify

arXiv · AI, language, vision and robotics · article · Sep 8, 2026 · UTC

Repeated-sampling evaluations increasingly extrapolate pass@k far beyond the number n of samples collected per problem. We show that, in the pooled/random-task conditional-Binomial model, fixed-n success counts identify only the n free moments of the latent per-task success distribution. Consequently, direct pass@k is identified for k n, even with arbitrarily many exchangeable tasks at the same rollout budget. This is stronger than the observation that the usual estimator is undefined beyond n: it characterizes the information missing from the fixed-depth count-law experiment. We give exact co

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:22:01.598Z. This is not the publication date.