AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts

arXiv · AI, language, vision and robotics · article · Sep 22, 2026 · UTC

A higher evaluation score does not always mean a better language model system. When optimization exploits an evaluator's mistakes, measured progress can conceal unchanged or deteriorating task performance. This failure can arise through parameter updates, selection among generated outputs, or revisions to persistent prompts. We develop a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text. Building on the Proxy Compression Hypothesis and research on inference-time and in-context reward hacking, we examine how reachable behavior, opt

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T04:21:13.910Z. This is not the publication date.