AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Beyond Correctness: Validity-Oriented Evaluation of Biomedical LLM Judges

arXiv · AI, language, vision and robotics · article · Aug 29, 2026 · UTC

We propose a scalable, validity-oriented pipeline for evaluating biomedical LLM judges when high-quality human judgments are scarce. First, we augment existing human-labelled biomedical benchmarks with deterministic, metric-grounded mutations that produce auditable preference pairs. Second, we evaluate judges beyond aggregate correctness using three deployment-relevant dimensions: correctness against metric-derived gold labels, robustness under repeated stochastic sampling, and compliance with the requested output format. We use this pipeline to assess Llama-3.1-8B-Instruct under four regimes:

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:51:58.603Z. This is not the publication date.