AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Commit-first LLM judging inherits the judge's own errors

arXiv · AI, language, vision and robotics · article · Aug 31, 2026 · UTC

LLM judges, models that score another system's output, can be gamed by the systems they score. Recent work identifies one defence that works: the judge solves the task itself first and commits to that answer, then accepts a candidate only if the two match. We call this commit-first judging, and ask whether shipped software implements it, and what it costs. We audit the default judge configurations of eight widely used evaluation frameworks. Of the 24 configurations in scope, none implement it. Nine implement a variant the literature measures as ineffective, and share one ancestor prompt, trace

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:01:58.596Z. This is not the publication date.