AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Prompt Breadth and Rollout Refresh Interact in On-Policy Distillation

arXiv · AI, language, vision and robotics · article · Sep 6, 2026 · UTC

How many prompts does on-policy distillation (OPD) need, and how does the answer depend on the student policies that generate its training responses? We study these two controls jointly: prompt breadth and rollout refresh. A 3x3 mathematical-reasoning experiment fixes 14,080 trajectories and 110 optimizer updates while varying the prompt bank and the number of response-generating policy snapshots. With ten snapshots, eight prompts reach 24.09% average accuracy, close to 24.51% for 14,080 distinct prompts. With responses frozen at the initial policy, however, increasing breadth lowers accuracy

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-25T16:52:32.424Z. This is not the publication date.