AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Matryoshka attribution: Learning to attribute language model outputs to representations and weights

arXiv · AI, language, vision and robotics · article · Sep 22, 2026 · UTC

Attributing language model outputs to their internal computations is an open problem in interpretability. Existing methods, which use causal interventions, gradients, or learnable masks, either are infeasibly expensive or struggle to identify actual causally-important internal computations. We propose framing attribution as the problem of identifying nested subsets of internal components which minimise a downstream loss. To learn this task, we introduce Matryoshka Attribution (MAttr), a mask learning method that parametrises the mask with a simple differentiable sigmoid top-$k$ operator. We su

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T04:21:13.910Z. This is not the publication date.