AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

From Trait Vectors to Circuits: Tracing Refusal and Sycophancy Through Language Models

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

A direction in activation space that changes safety-relevant behavior when steered is not necessarily one the model uses to produce that behavior on its own. We therefore ask whether steering acts through the computation of the unmodified model or through a different set of components, studying two traits whose directions have been extracted and validated in prior work: refusal and sycophancy in Qwen2.5-7B-Instruct. For each, we use the trait vector to split the computation into a reconstruction circuit before the vector and a transmission circuit after it. We then test whether restoring the c

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-26T06:21:50.202Z. This is not the publication date.