SOURCE-LINKED INTELLIGENCE
VoiceCodeBench: Evaluating Exact Structured-Token Recovery in Automatic Speech Recognition
Automatic speech recognition is usually evaluated with word error rate (WER), although voice workflows often require exact written values. VoiceCodeBench measures whether transcripts preserve identifiers, paths, commands, and other structured tokens needed by downstream software. It contains 300 human-recorded English workplace segments (5.59 hours, 85 speakers) and 1,482 audited entities across 26 types and eight domains. Under a raw-audio-only protocol, we evaluate 19 batch and streaming systems using WER, Canonical Token/Entity Match (CTEM), and strict segment-level Task Success Rate (TSR).
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-08-28T22:23:42.000Z
First collected: 2026-09-21T08:02:06.831Z. This is not the publication date.