AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

KerColle: Unlocking Fine-Grained GPU Concurrency in Vision-Language-Action Models

arXiv · AI, language, vision and robotics · article · Sep 16, 2026 · UTC

Vision-Language-Action (VLA) models have emerged as foundational models for next-generation robotics. High VLA inference throughput is critical for meeting the control-rate requirements of robots. VLA models comprise two phases, a vision-language model (VLM) and an action head, that can be decoupled and executed asynchronously and concurrently across independent robot requests. Through a detailed characterization of four state-of-the-art VLAs, we observe that GPUs are severely underutilized in VLA inference as the Cooperative Thread Array (CTA) scheduler of GPU is unable to fully overlap the t

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T17:51:24.264Z. This is not the publication date.