cd /news/large-language-models/latentport-beyond-kv-cache-cross-mod… · home topics large-language-models article
[ARTICLE · art-137799] src=arxiv.org ↗ pub= topic=large-language-models verified=true sentiment=· neutral

LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay

Researchers demonstrated the first cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay, transferring live memory from a Qwen3.5 4B model to a 9B sibling, according to an arXiv paper (2609.25053v1). Adding the Gated DeltaNet (GDN) persistent-state package to translated attention KV lowered teacher-forced negative log-likelihood by 0.747 nats/token (95% paired document bootstrap CI [0.6921, 0.8047]), improving all 64 PG19 documents, and a 434,176-parameter correction brought continuation loss to 0.076 nats/token above native 9B with Jensen-Shannon divergence of 0.022 and native context recovery of 0.918. The authors state the evidence covers one direction, one geometry-matched Base-model pair, and 4K teacher-forced continuation, with the near-native gate failed, the 16K branch not run, and free-generation equivalence and a general state interface unproven.

by read1 min views1 publishedSep 23, 2026

arXiv:2609.25053v1 Announce Type: new Abstract: Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our knowledge, this is the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay. Translated attention KV alone leaves a large gap; adding the Gated DeltaNet (GDN) persistent-state package lowers teacher-forced negative log-likelihood (NLL), the average next-token log-loss, by 0.747 nats/token (95% paired document bootstrap CI [0.6921, 0.8047]), improving all 64 PG19 documents. Direct recurrent and convolution reuse outperforms the tested learned GDN maps, consistent with partial functional compatibility of persistent-state coordinates. A fresh component factorial selects translated KV with direct recurrent and convolution state. An additional 434,176-parameter correction improves that base on 64 fresh web documents: continuation loss is 0.076 nats/token above native 9B (excess NLL), Jensen-Shannon (JS) divergence is 0.022, and native context recovery (NCR) is 0.918. Corrected 9B significantly beats continued 4B inference while processing zero historical prefix tokens. Evidence covers one direction, one geometry-matched Base-model pair, and 4K teacher-forced continuation; the near-native gate failed, the 16K branch was not run, and free-generation equivalence and a general state interface remain unproven.

── more in #large-language-models 4 stories · sorted by recency
── more on @qwen3.5 4b 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/latentport-beyond-kv…] indexed:0 read:1min 2026-09-23 ·