AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

arXiv · AI, language, vision and robotics · article · Sep 1, 2026 · UTC

Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. A teacher biased by a system prompt can generate semantically clean training data, such as numeric sequences, that still causes a downstream student to inherit the hidden preference, a phenomenon known as subliminal learning. Prior work has identified several parts of this process. How the signal builds up during training and produces behavioral transfer remains unclear, making targeted mitigation difficult. We propose and validate trait-direction drift as a mechanism for subliminal learning: biased gen

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T06:11:57.537Z. This is not the publication date.