AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Towards Zero-Shot Transfer Across Embodiments For Driving VLAs

arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC

Vision-Language-Action models (VLAs) have shown strong potential in autonomous driving by leveraging multimodal pretraining for instruction following, visual reasoning, and scene-level generalization. In robotic manipulation, scaling VLA fine-tuning across multiple robot setups--especially when unifying representations across embodiments--has been shown to improve in-dataset performance and cross-embodiment generalization; in autonomous driving, however, VLAs remain largely trained on individual datasets and are rarely evaluated for zero-shot transfer to unseen datasets and camera rigs; furthe

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T05:32:15.665Z. This is not the publication date.