{"slug": "task-oriented-key-layer-kv-communication-for-efficient-latent-multi-agent", "title": "Task-Oriented Key-Layer KV Communication for Efficient Latent Multi-Agent Collaboration", "summary": "Researchers proposed KITE, a training-free framework for task-oriented key-layer KV communication in latent multi-agent LLM collaboration, detailed in arXiv paper 2610.08820v1. KITE selects a task-effective key layer via a receiver trajectory distortion criterion, transmits only that layer's latent working memory, and uses the same layer as the entry point for autoregressive latent reasoning. Across seven benchmarks, two model families, and three model scales, KITE cut communication volume by 28-36x, delivered up to 3x end-to-end inference speedup, and improved accuracy by up to 23.3 percentage points versus full-layer KV communication.", "body_md": "arXiv:2610.08820v1 Announce Type: new \nAbstract: Large language model-based multi-agent systems improve complex problem solving through collaboration, while latent communication directly transmits model internal states to avoid the high inference costs of natural language. However, existing KV-based latent communication methods prioritize sender-side state fidelity, leading to substantial communication and computation overhead and potentially introducing redundant information. To address these limitations, we revisit latent communication from a task-oriented perspective, shifting its objective from sender-side state fidelity to receiver-side task sufficiency. Under this formulation, we propose KITE, a training-free framework for task-oriented key-layer KV communication. KITE identifies a task-effective key layer using a receiver trajectory distortion criterion, transmits only the latent working memory associated with the key layer, and further uses the same layer as the entry point for autoregressive latent reasoning. Experiments on seven benchmarks across two model families and three model scales show that, compared with full-layer KV communication, KITE reduces communication volume by 28-36$\\times$, achieves up to 3$\\times$ end-to-end inference speedup, and improves accuracy by up to 23.3 percentage points.", "url": "https://wpnews.pro/news/task-oriented-key-layer-kv-communication-for-efficient-latent-multi-agent", "canonical_source": "https://arxiv.org/abs/2610.08820", "published_at": "2026-10-08 04:00:00+00:00", "updated_at": "2026-10-08 04:18:58.723647+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-research", "machine-learning"], "entities": ["KITE", "arXiv", "2610.08820v1"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/task-oriented-key-layer-kv-communication-for-efficient-latent-multi-agent", "markdown": "https://wpnews.pro/news/task-oriented-key-layer-kv-communication-for-efficient-latent-multi-agent.md", "text": "https://wpnews.pro/news/task-oriented-key-layer-kv-communication-for-efficient-latent-multi-agent.txt", "jsonld": "https://wpnews.pro/news/task-oriented-key-layer-kv-communication-for-efficient-latent-multi-agent.jsonld"}}