SOURCE-LINKED INTELLIGENCE
Are Near-Tied LLM Rankings Robust to Family-DIF-Guided Benchmark Recomposition?
Small leaderboard gaps are often interpreted as evidence that one language model is better than another, but their sign may depend on which benchmark items are included. We test this using item-level responses from five benchmarks and a family-label-free spectral approximation to multidimensional item-response theory (MIRT). In owner-disjoint folds, one owner half identifies items with low residual differential item functioning across model families (low-DIF); the resulting frozen, source- and easiness-balanced weights score models in the other half, while equally short matched-random subtests
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-08-31T23:29:50.000Z
First collected: 2026-09-21T06:21:59.299Z. This is not the publication date.