AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Scaling Vision-Language Reward Learning for Robot Manipulation in Parallel Simulation

arXiv · AI, language, vision and robotics · article · Sep 18, 2026 · UTC

Vision-language models (VLMs) can replace human annotators in preference-based reward learning, but sequential API requests and single-environment data collection make training slow and costly. We present RAPID (Reward learning with Adaptive Parallel Image Diversity), a system that couples GPU-parallel rollout with data-aware policy updates, single-request preference labeling, automatic reward stabilization, and representative image sampling. We evaluate these components on five Franka Panda manipulation tasks in IsaacLab. Parallel rollout and adaptive updates provide the first substantial red

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T13:51:27.104Z. This is not the publication date.