cd /news/artificial-intelligence/voice-memory-for-agentic-speech-reco… · home topics artificial-intelligence article
[ARTICLE · art-79697] src=machinebrief.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Voice Memory for Agentic Speech Recognition

Voice Memory, a new inference-only scheme for agentic speech recognition, reduces weighted word error rate from 8.36% to 7.52% across ten HyPoradise domains without regressing any dataset below its 1-best baseline, according to a paper on arXiv. The method uses a frozen corrector and a score-gated optimizer that revises a per-domain memory file through bounded edits, cutting over-correction errors from 64% to 35% on financial news. Gains are largest in air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%), and the memory transfers across corrector families with zero added parameters.

read1 min views1 publishedJul 30, 2026

arXiv:2607.26410v1 Announce Type: new Abstract: We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain memory.md and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The memory transfers across corrector families and adds zero parameters to the inference path. A demo and example code are provided for future studies.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @voice memory 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/voice-memory-for-age…] indexed:0 read:1min 2026-07-30 ·