SOURCE-LINKED INTELLIGENCE
Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing
Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works address this problem by conditioning the speech synthesis process on lip movements extracted from the video signal. In this work, we condition the speech generation on a binary voice-activity signal, which has a lightweight representation and can be produced in multiple ways. We show that the model follows the voice-activity signal with high accuracy w
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-22T14:24:10.000Z
First collected: 2026-09-23T04:11:12.117Z. This is not the publication date.