SOURCE-LINKED INTELLIGENCE
EPIG-Tree: Compute-Optimal Branching for Gradient-Efficient Reinforcement Learning
Reward-based reinforcement learning for language models, exemplified by Group Relative Policy Optimization (GRPO), collapses an entire stochastic trajectory into a single scalar reward. This is clean and scalable, but it explores and allocates reward inefficiently: a trajectory may contain many causal decisions, recovery attempts, and environment-randomness events, yet every token or action inherits one trajectory-level advantage. We study tree-based rollout construction as a compute-allocation problem for policy-gradient estimation. Our central claim is that branches should be placed not wher
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T10:08:48.000Z
- arXiv · Artificial Intelligence · 2026-09-17T10:08:48.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.