SOURCE-LINKED INTELLIGENCE
I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance
Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearable applications for Deaf and Hard of Hearing users. We propose Interrupt and Silent Modeling (ISM), a model-agnostic paradigm that embeds proactive decisions into LLM decoding via two special tokens: \texttt{ } and \texttt{ }, capturing four states: onset detection, sustained-relevance triggering, irrelevance suppressi
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-18T01:04:53.000Z
First collected: 2026-09-23T14:01:59.594Z. This is not the publication date.