AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms

arXiv · AI, language, vision and robotics · article · Sep 20, 2026 · UTC

Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While existing solutions rely on external priors or specialized data, we investigate the root cause by exploring the internal mechanisms of these models. Specifically, we present the first interpretability study on the ''motion planning'' process of text-to-video diffusion models, revealing how motion trajectories form during early denoising stages. Building upon the ''first shape, then details'' finding, we combine cross-attention trajectory patterns with c

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T09:51:33.063Z. This is not the publication date.