AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Reference-Grafting Matches Fine-Tuning at Eliciting Sandbagged Capabilities

arXiv · AI, language, vision and robotics · article · Aug 29, 2026 · UTC

Sandbagging, in which a model deliberately underperforms on an evaluation despite retaining the underlying capability, threatens the safety evaluations that frontier-model governance depends on. The Elicitation Game found that fine-tuning elicits hidden capability from sandbagging model organisms whereas additive activation steering fails. We revisit that verdict with reference-grafting, which sets an activation's coordinate along a contrast direction to the value it takes in an honest reference, at a small set of circuits chosen by active learning. Across eleven fine-tuned password-locked mod

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:31:56.984Z. This is not the publication date.