{"slug": "compressing-streaming-neural-audio-encoders-via-latent-space-distillation", "title": "Compressing Streaming Neural Audio Encoders via Latent-Space Distillation", "summary": "Apple researchers compressed an on-device streaming neural audio tokenizer by 2.8× using latent-space distillation, keeping the distilled student within 1.9% relative word error rate of its teacher on five of six teacher–student pairs without fine-tuning. The method trains only the student encoder to regress the teacher's pre-quantizer latent under a squared-error objective, with a single affine layer absorbing the teacher–student width mismatch, and beats an independently trained tokenizer of identical capacity by 3.9% relative. The work targets the always-on tokenizer in System-wide Dictation on Apple devices, where the encoder competes for DRAM with a sparsely activated, Instruction-Following-Pruned foundation model.", "body_md": "System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes—the last representation the two token interfaces share. We train only the student encoder to regress the teacher’s per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher–student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8× compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher–student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.\n\n- ‡ Equal contribution\n- † NVIDIA\n- § Anthropic\n- ** Work done while at Apple", "url": "https://wpnews.pro/news/compressing-streaming-neural-audio-encoders-via-latent-space-distillation", "canonical_source": "https://machinelearning.apple.com/research/latent-space-distillation", "published_at": "2026-09-24 00:00:00+00:00", "updated_at": "2026-09-24 15:59:55.474284+00:00", "lang": "en", "topics": ["artificial-intelligence", "machine-learning", "natural-language-processing", "ai-research"], "entities": ["Apple", "NVIDIA", "Anthropic", "System-wide Dictation", "Instruction-Following Pruning"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/compressing-streaming-neural-audio-encoders-via-latent-space-distillation", "markdown": "https://wpnews.pro/news/compressing-streaming-neural-audio-encoders-via-latent-space-distillation.md", "text": "https://wpnews.pro/news/compressing-streaming-neural-audio-encoders-via-latent-space-distillation.txt", "jsonld": "https://wpnews.pro/news/compressing-streaming-neural-audio-encoders-via-latent-space-distillation.jsonld"}}