{"slug": "latentport-beyond-kv-cache-cross-model-transfer-of-recurrent-memory-in-hybrid-a", "title": "LatentPort: Beyond KV Cache - Cross-Model Transfer of Recurrent Memory in Hybrid Language Models: A 4B-to-9B Hybrid-State Handoff Without Target Prefix Replay", "summary": "Researchers demonstrated the first cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay, transferring live memory from a Qwen3.5 4B model to a 9B sibling, according to an arXiv paper (2609.25053v1). Adding the Gated DeltaNet (GDN) persistent-state package to translated attention KV lowered teacher-forced negative log-likelihood by 0.747 nats/token (95% paired document bootstrap CI [0.6921, 0.8047]), improving all 64 PG19 documents, and a 434,176-parameter correction brought continuation loss to 0.076 nats/token above native 9B with Jensen-Shannon divergence of 0.022 and native context recovery of 0.918. The authors state the evidence covers one direction, one geometry-matched Base-model pair, and 4K teacher-forced continuation, with the near-native gate failed, the 16K branch not run, and free-generation equivalence and a general state interface unproven.", "body_md": "arXiv:2609.25053v1 Announce Type: new \nAbstract: Can one language model hand its live memory to another without the receiver rereading the context? We demonstrate useful persistent hybrid-state transfer across one architecture-matched Qwen3.5 4B-to-9B sibling pair. To our knowledge, this is the first demonstrated cross-model handoff of persistent recurrent inference state between differently sized hybrid language models without target prefix replay. Translated attention KV alone leaves a large gap; adding the Gated DeltaNet (GDN) persistent-state package lowers teacher-forced negative log-likelihood (NLL), the average next-token log-loss, by 0.747 nats/token (95% paired document bootstrap CI [0.6921, 0.8047]), improving all 64 PG19 documents. Direct recurrent and convolution reuse outperforms the tested learned GDN maps, consistent with partial functional compatibility of persistent-state coordinates. A fresh component factorial selects translated KV with direct recurrent and convolution state. An additional 434,176-parameter correction improves that base on 64 fresh web documents: continuation loss is 0.076 nats/token above native 9B (excess NLL), Jensen-Shannon (JS) divergence is 0.022, and native context recovery (NCR) is 0.918. Corrected 9B significantly beats continued 4B inference while processing zero historical prefix tokens. Evidence covers one direction, one geometry-matched Base-model pair, and 4K teacher-forced continuation; the near-native gate failed, the 16K branch was not run, and free-generation equivalence and a general state interface remain unproven.", "url": "https://wpnews.pro/news/latentport-beyond-kv-cache-cross-model-transfer-of-recurrent-memory-in-hybrid-a", "canonical_source": "https://arxiv.org/abs/2609.25053", "published_at": "2026-09-23 04:00:00+00:00", "updated_at": "2026-09-23 04:25:36.331629+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "machine-learning", "natural-language-processing"], "entities": ["Qwen3.5 4B", "Qwen3.5 9B", "Gated DeltaNet", "PG19", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/latentport-beyond-kv-cache-cross-model-transfer-of-recurrent-memory-in-hybrid-a", "markdown": "https://wpnews.pro/news/latentport-beyond-kv-cache-cross-model-transfer-of-recurrent-memory-in-hybrid-a.md", "text": "https://wpnews.pro/news/latentport-beyond-kv-cache-cross-model-transfer-of-recurrent-memory-in-hybrid-a.txt", "jsonld": "https://wpnews.pro/news/latentport-beyond-kv-cache-cross-model-transfer-of-recurrent-memory-in-hybrid-a.jsonld"}}