AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Enriching Speech Emotion Representations with Conversational Context

arXiv · AI, language, vision and robotics · article · Sep 22, 2026 · UTC

Detecting emotions is necessary for building systems that can accurately and adaptively interact with humans. Speech Emotion Recognition (SER) has become an important research focus to develop intelligent spoken interfaces. However, most studies predict emotions at the utterance level, ignoring the conversational context, along with the emotional flow and speaker interactions it carries. In this paper, we introduce ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T04:11:12.117Z. This is not the publication date.