AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing

arXiv · AI, language, vision and robotics · article · Sep 22, 2026 · UTC

Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works address this problem by conditioning the speech synthesis process on lip movements extracted from the video signal. In this work, we condition the speech generation on a binary voice-activity signal, which has a lightweight representation and can be produced in multiple ways. We show that the model follows the voice-activity signal with high accuracy w

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T04:11:12.117Z. This is not the publication date.