cd /news/artificial-intelligence/nexus-depth-adaptive-kv-cache-splici… · home topics artificial-intelligence article
[ARTICLE · art-108228] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

Nexus: Depth-Adaptive KV-Cache Splicing and Retrieval-Decoupled Tool Routing for Agentic LLMs on Unified Memory

Nexus, a depth-adaptive KV-cache splicing and retrieval-decoupled tool routing system for agentic LLMs on unified memory, reduces time-to-first-token by up to 1.7x and saves ~80% of main-context tokens by using an INT8 semantic lookaside buffer with a calibrated cross-encoder margin gate, achieving near 89% routing accuracy with 250 tools. The system, tested on Qwen2.5-14B-Instruct Q4_K_M on Apple silicon, guarantees output fidelity (top-1 agreement, D_KL ≈ 0) but notes latency can dip to 0.98x before converging to parity, and reports two negative results: off-anchor RoPE fidelity boundary and a reference-free drift gate failure (Spearman rho = 0.193).

read1 min views1 publishedAug 24, 2026

arXiv:2608.20397v1 Announce Type: new Abstract: Agentic large language models (LLMs) on the Model Context Protocol (MCP) re-encode verbose tool schemas every turn, so prefill - quadratic in sequence length - dominates time-to-first-token (TTFT) as the tool registry grows. Nexus's primary lever is to decouple routing from the schema-prefill cost: an INT8 semantic lookaside buffer (SLB) with a calibrated cross-encoder margin gate selects tools by retrieval, and arguments are generated over a compressed textual signature (median 19 tokens) rather than over spliced key/value (KV) cache. This path is depth-independent: routing accuracy stays near 89% as the registry scales to 250 tools - where a concatenate-all-schemas baseline overflows the context window entirely - and it reaches a first-argument token 1.66x sooner than a full-schema re-prefill at a ~80% main-context token saving. As a secondary, bounded lever we transplant a compiled schema KV block directly into the live context. This is fundamentally limited by rotary position embedding (RoPE) phase drift: an anchored splice is output-exact, but off-anchor placement corrupts attention, so beyond a threshold P=256 Nexus repairs the seam with a depth-adaptive suffix redecode that escalates to a full re-prefill. The resulting never-regress property is a guarantee on output fidelity (top-1 agreement, D_KL approx. 0) - not on latency, which can dip to 0.98x before converging to parity - alongside a 1.1-1.7x TTFT speedup at moderate depth that narrows to parity at deep context. Two negative results bound the design: the off-anchor RoPE fidelity boundary, and the failure of a reference-free drift gate to predict drift (Spearman rho = 0.193). All measurements are from one model tuple (Qwen2.5-14B-Instruct Q4_K_M) on Apple-silicon unified memory; the qualitative boundaries generalize, while the quantitative envelope is tuple-specific.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @nexus 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/nexus-depth-adaptive…] indexed:0 read:1min 2026-08-24 ·