AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Modeling, Scaling, and Decoding: Optimizing Controllable Speech Generation with Nonverbal Vocalizations

arXiv · AI, language, vision and robotics · article · Sep 13, 2026 · UTC

Controllable synthesis of nonverbal vocalizations (NVVs) is es- sential for natural and expressive speech, but remains challeng- ing due to their acoustic diversity and imbalanced distribution in existing corpora. To address these challenges, we develop an NVV-aware DiTAR system that models continuous speech latents, encodes the 16 target NVV categories as dedicated to- kens, and adapts stop prediction to distinguish mid-utterance vocalizations from utterance boundaries. Training begins with large-scale bilingual pre-training on diverse NVV speech, fol- lowed by continued supervised fine-tunin

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T12:41:04.663Z. This is not the publication date.