SOURCE-LINKED INTELLIGENCE
Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
A higher evaluation score does not always mean a better language model system. When optimization exploits an evaluator's mistakes, measured progress can conceal unchanged or deteriorating task performance. This failure can arise through parameter updates, selection among generated outputs, or revisions to persistent prompts. We develop a comparative framework for reward hacking across these three optimization substrates: weights, selection, and text. Building on the Proxy Compression Hypothesis and research on inference-time and in-context reward hacking, we examine how reachable behavior, opt
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-22T08:12:58.000Z
First collected: 2026-09-23T04:21:13.910Z. This is not the publication date.