AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Auditing Generative Audio Calls for Known-Task Audio-LLM Evaluatio

arXiv · AI, language, vision and robotics · article · Aug 28, 2026 · UTC

Speech and audio LLMs are often evaluated by asking whether a waveform prompt beats an automatic speech recognition (ASR) transcript. For known closed-set tasks, that comparison conflates two factors: access to acoustic evidence and the need to call a generative audio model. We evaluate this distinction as a controlled call-decision problem. For each example, a policy chooses among keeping a transcript label, using encoder evidence from Contrastive Language-Audio Pretraining (CLAP), Audio Spectrogram Transformer (AST), or WavLM, and calling Qwen2-Audio, Qwen2.5-Omni, or MOSS-Audio; the decisiv

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T08:21:55.975Z. This is not the publication date.