AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they claim to measure. When presented with a terse correct solution and a commented buggy solution for the same MBPP problem, a public preference reward model selects the correct one no better than a coin flip (0.507). Subtracting the predictable surface component from such scores is increasingly common, but removal alone does not yield a more valid measurement: the removed component may carry construct-relevant signal, and residualization cannot tell

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T08:01:43.213Z. This is not the publication date.