AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Multimodal Voice Activity Projection for Social Robot Mediation: Expected Behavior and Deployment Constraints

arXiv · AI, language, vision and robotics · article · Sep 23, 2026 · UTC

Turn-taking prediction is especially relevant for social robots that act as mediators in human-human interaction, where the expected action is often not to speak, but to orient, wait, avoid interruption, or prepare a balanced intervention. This paper presents Multimodal Voice Activity Projection (MM-VAP) as a human state-aware perception layer for future robot mediation behavior. The model estimates the future evolution of the conversational floor from synchronized audio-visual evidence and derives turn-taking events such as Hold, Shift, Shift prediction, Backchannel prediction, and overlap-re

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-24T08:22:30.429Z. This is not the publication date.