SOURCE-LINKED INTELLIGENCE
Compressing Streaming Neural Audio Encoders via Latent-Space Distillation
System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision…
Read original source ↗ Open in workspace
- recordType
- article
- region
- Global
Evidence & attribution
- Apple Machine Learning Research · 2026-09-24T00:00:00.000Z
First collected: 2026-09-24T17:52:48.368Z. This is not the publication date.