AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Three Ways Classical Test Theory Misleads for LLM Judges

arXiv · AI, language, vision and robotics · article · Aug 31, 2026 · UTC

An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrowing reliability statistics from classical test theory, usually without stating the measurement design each statistic assumes, and we show that three widely portable ones mean something different for a judge than for a test because the judge setting rearranges the roles those designs rest on. First, an internal-consistency coefficient computed over rubric elements contains no scorer facet. Holding one judge's measured error rate fixed at $4

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-26T17:51:55.454Z. This is not the publication date.