SOURCE-LINKED INTELLIGENCE
Stable and Efficient Real-World Online VLA Post-Training via Asynchronous Replay-Anchored Policy Improvement
Online post-training of vision-language-action (VLA) models requires efficient use of robot interaction and reliable policy improvement from continually collected experience. We propose asynchronous Replay-Anchored Policy improvement (RAPolicy), a framework that performs rollout and learning concurrently while grounding both critic and actor updates in replayed behavior. The critic learns chunk-level values from recorded actions and constructs Bellman targets without predicting next actions, reducing computation and dependence on action-value estimates outside replay coverage. The one-step flo
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-19T08:45:18.000Z
First collected: 2026-09-23T12:01:45.602Z. This is not the publication date.