AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Look Before You Steer: Geometry Predicts SAE Feature Steerability

arXiv · AI, language, vision and robotics · article · Sep 19, 2026 · UTC

Steering with SAE features requires per-feature coefficient tuning, which currently demands intervention sweeps. We ask whether properties of the SAE itself, computable before any forward pass, predict which features will be cheap or expensive to steer. We show that variation in SAE feature steerability is partially predicted by decoder-space geometry: neighbor density and maximum cosine similarity to nearby decoder directions, both computable from the SAE weight matrix before any intervention, rank features by how much steering they require for a fixed behavioral effect ($ρ$ up to $-0.546$, $

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T12:01:45.602Z. This is not the publication date.