AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

arXiv · AI, language, vision and robotics · article · Sep 18, 2026 · UTC

Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user's underlying demand from complex multimodal interaction? Real-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history. This inference is further complicated by

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-23T14:01:59.594Z. This is not the publication date.