AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Luck Is Not Skill: When Do Paired Rollouts Help Group-Relative RL of LLM Agents?

arXiv · AI, language, vision and robotics · article · Sep 21, 2026 · UTC

Group-relative reinforcement learning compares rollouts of the same prompt, but independent environment noise can obscure these comparisons. We study paired rollouts, which share an event-keyed noise schedule within each group while preserving each rollout's marginal distribution. Pairing removes the between-schedule component of reward-contrast variance, but need not reduce gradient variance. For one-sided grader noise, we derive an exact condition for reduction and give a counterexample in which reward contrasts improve while gradient variance increases. A controlled study trains a 2B tool-u

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T08:01:43.213Z. This is not the publication date.