AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

On Mitigation of Subliminal Learning in Large Language Models

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

Knowledge distillation can transmit unintended behavioral traits from a teacher model to a student through training data that appear semantically unrelated to those traits, a phenomenon known as subliminal learning. Although recent work has established this effect, its training dynamics and mitigation remain underexplored. We study subliminal learning in open-weight language models ranging from 1.5B to 8B parameters, covering the Qwen, Gemma, and Llama families in number-sequence and chain-of-thought settings. Rather than evaluating only final models, we track trait-related probabilities throu

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-26T08:21:45.852Z. This is not the publication date.