AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

A Causal Model for Locating and Unlocking Sandbagging in Model Organisms

arXiv · AI, language, vision and robotics · article · Aug 29, 2026 · UTC

Sandbagging models strategically underperform on evaluations while retaining the capabilities being measured. The evaluations that guide frontier-model deployment and governance then understate what these models can do. To understand the mechanism, we propose a causal model of how sandbagging is carried in the residual stream. Early layers write the sandbagging intent onto a single axis of the stream, and a later layer reads that axis and commits the answer. We study three instruction-tuned models (Qwen2.5-7B, Llama-3-8B, and Mistral-7B) and four ways of installing a sandbagging lock (promptin

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T07:31:56.984Z. This is not the publication date.