AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

The Attention Triangle in Audio-Video Models

arXiv · AI, language, vision and robotics · article · Sep 3, 2026 · UTC

Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This e

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T04:51:57.792Z. This is not the publication date.