AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

OPDSearch+: On-Policy Distillation with RL Refinement for Search-Augmented Reasoning

arXiv · AI, language, vision and robotics · article · Aug 25, 2026 · UTC

Search-augmented reasoning remains difficult for small language models. On-policy distillation (OPD) from trained teachers offers a promising direction, but suffers from two issues: (1) high-quality multi-turn search trajectories depend on dynamic retriever responses, making SFT data prohibitively expensive to collect at scale; (2) task-specifically trained teachers incur substantial training cost, while directly applying OPD with an off-the-shelf teacher without task-specific fine-tuning constrains the student to the teacher's performance ceiling and suffers from severe training instability.

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T10:02:02.728Z. This is not the publication date.