AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

SteerCheck: Attribution Specificity and Alignment Leakage in Activation-Steering Audits

arXiv · AI, language, vision and robotics · article · Aug 25, 2026 · UTC

Activation steering can change behaviour without establishing that the effect is specific to the intended concept. We introduce SteerCheck, a preregistered attribution audit that matches off-target KL and separates mean, protected-tail, polarity, transfer, and semantic claims. Exact replay of 960 Qwen3-14B interventions reveals complementary limits of common controls: isotropic directions occupy a narrow near-orthogonal region, whereas sign-randomized same-construction directions often retain substantial target alignment. Effect is strongly associated with signed cosine within the sign-randomi

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T10:02:02.728Z. This is not the publication date.