AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

What Matters, When? Diagnosing and Improving Conditional Visual Grounding in Visuomotor Imitation Policies

arXiv · AI, language, vision and robotics · article · Sep 4, 2026 · UTC

Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that dist

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T21:52:07.471Z. This is not the publication date.