AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

VISTA: Video-Injected Stylized Text-to-Animation

arXiv · AI, language, vision and robotics · article · Sep 20, 2026 · UTC

We present VISTA, a two-stage framework for generating stylized 3D human motion by fusing structural content from text prompts with expressive style from reference videos, without requiring jointly paired (text, video, stylized motion) triplets. A Dual-channel Autoencoder first maps motion sequences and video clips into a shared latent manifold. A masked autoregressive diffusion backbone then operates within this manifold, injecting video-derived style through a dedicated late-fusion Dual-AdaLN pathway while preserving text-conditioned content structure. A cross-batch unpaired training protoco

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T09:51:33.063Z. This is not the publication date.