SOURCE-LINKED INTELLIGENCE
Shared Actors Need Not Share Critics: Effects of Value Mismatch in Parallel Reinforcement Learning
When a single policy is trained in parallel across multiple environments of the same task, such as procedurally generated levels, randomized dynamics, or curricula, implementations commonly use one critic across all sampled environments. Yet different environments can assign different expected returns to the same input visible to the critic. A critic without environment information must then reconcile distinct value targets, systematically shifting the sampled advantages within individual environments. Using illustrative bandit models with multiple environments and a common optimal arm, we cha
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-08-27T00:12:15.000Z
First collected: 2026-09-21T09:11:58.312Z. This is not the publication date.