SOURCE-LINKED INTELLIGENCE
VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis
arXiv · AI, language, vision and robotics · article · Sep 2, 2026 · UTC
Speech systems increasingly infer how an utterance should be delivered from context, but a plausible delivery plan may not be supported by the input. VoxReason is a small public benchmark and verifier for testing this failure before waveform synthesis. Each of its 100 cases fixes the utterance, provides derived records that name the source emotion and intensity, and changes one licensed cue. A system must cite the record for its delivery decision and update only the plan fields associated with the edit. We use deterministic verifier references, not a model leaderboard. This holdout excludes ev
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
First collected: 2026-09-21T05:11:56.580Z. This is not the publication date.
Observed changes
AIIC observation times, not verified publisher revision times. Up to eight recent revisions.
2026-09-26T06:21:50.202Z
- title:
VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis → VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis - summary:
Expressive speech systems have to decide how an utterance is delivered before any waveform is rendered. In dialogue agents, narration, and role-conditioned TTS, that planning step sets affect, pitch, energy, rate, pause, emphasis, and stance, yet standard audio metrics rarely show whether those choices were actually licensed by the source record. This leaves a practical evaluation gap: a system may sound plausible while relying on a memorized script instead of the cue that governs delivery. VoxReason casts this pre-synthesis step as a listener-free task for source-grounded speech planning. Sys → Speech systems increasingly infer how an utterance should be delivered from context, but a plausible delivery plan may not be supported by the input. VoxReason is a small public benchmark and verifier for testing this failure before waveform synthesis. Each of its 100 cases fixes the utterance, provides derived records that name the source emotion and intensity, and changes one licensed cue. A system must cite the record for its delivery decision and update only the plan fields associated with the edit. We use deterministic verifier references, not a model leaderboard. This holdout excludes ev