AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

What Moves? Localized Motion Representations for Compositional Scene Control

arXiv · AI, language, vision and robotics · article · Sep 3, 2026 · UTC

Real-world dynamics are inherently compositional: multiple entities move simultaneously within a shared scene, each exhibiting distinct motion patterns. Yet current motion representation models entangle the dynamics of different entities, without explicitly capturing localized motion for each individually. Crucially, motion is defined relative to a global reference frame, including camera motion and scene layout. However, localized embeddings are often computed from cropped images or obtained by masking features after encoding, discarding the context needed to interpret motion. To address this

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T04:31:57.454Z. This is not the publication date.