AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Teacher Should Think Ahead: Adaptive Continuations for Reliable On-Policy Distillation

arXiv · AI, language, vision and robotics · article · Sep 6, 2026 · UTC

On-policy distillation (OPD) is a promising approach for transferring knowledge between language models, where a student receives dense token-level supervision along its own generated trajectories. However, teacher supervision can be unreliable when conditioned on incomplete or low-quality student prefixes. We identify Teacher Uncertainty Contraction (TUC), a systematic phenomenon whereby the teacher's predictive uncertainty decreases as it continues from a student-generated prefix. We theoretically characterize this trade-off through a variance-bias decomposition of teacher-branch gradients,

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-25T16:52:32.424Z. This is not the publication date.