AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

The Price of Consistency: Exploiting Visual Anchors for Multimodal Jailbreaking in Video Generation

arXiv · AI, language, vision and robotics · article · Sep 7, 2026 · UTC

The rapid evolution of video generation has shifted the paradigm from pure text-driven to multi-conditional controllable generation, with reference images now widely adopted as conditional inputs to achieve superior spatiotemporal consistency. While these reference images serve as powerful visual anchors that significantly enhance controllability, their impact on safety remains largely unexplored. In this work, we reveal the visual anchoring effect: by enforcing consistency, the mechanism prevents the generated content from drifting away from the original harmful intent, thereby eliminating th

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:52:10.320Z. This is not the publication date.