AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

How Sensitive Are LLM Leaderboard Claims to Hidden Model Selection?

arXiv · AI, language, vision and robotics · article · Sep 23, 2026 · UTC

LLM leaderboard gains can reflect selection among privately evaluated model variants, yet neither the number of variants nor their dependence is public. We ask how many hidden variants a published margin can support while retaining statistical evidence of a provider's advantage over a fixed comparator. For a fixed candidate family under a Gaussian margin model, we derive a sensitivity curve that reports this maximum count as a function of a lower bound on within-family correlation. The relevant correlation must match the score used for ranking and the sampling model: in a controlled family, po

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-24T08:22:30.429Z. This is not the publication date.