AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Validating, Not Sampling: Region-Level Robustness of Vision-Language and Vision-Language-Action Models

arXiv · AI, language, vision and robotics · article · Sep 14, 2026 · UTC

Vision-language models (VLMs) and vision-language-action models (VLAs) are increasingly deployed in real-world applications. There, a small perturbation to the recorded camera image may change a decision significantly. However, existing benchmarks for these models only sample perturbations, which does not guarantee the absence of a failure in the untested region. We present the first robustness validation of six VLMs (drawn from the Gemma, InternVL, LLaVA, and Qwen families) and five VLAs (drawn from the GR00T, OpenVLA, and $π$ families) over entire continuous regions of photometric and geomet

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-24T08:32:17.684Z. This is not the publication date.