{"slug": "composable-cxl-memory-as-a-kubernetes-native-shared-memory-for-llm-serving", "title": "Composable CXL Memory as a Kubernetes-Native Shared Memory for LLM Serving", "summary": "A Kubernetes Dynamic Resource Allocation (DRA) driver makes composable CXL memory a schedulable cluster resource for cross-node KV-cache reuse in LLM serving, according to an arXiv paper (2609.10790v1). On a two-node cluster with a 512 GiB CXL appliance and Qwen2.5-7B-Instruct, cross-node prefix reuse cut time-to-first-token by 5.5x to 36.6x at an external hit rate of 95.4-99.5%, while node-local tiers (GPU prefix caching, CPU-DRAM offload) fell back to full recompute. The authors report the sharing gap, the latency ratio between cross-node and same-node reuse, at 1-4%, and describe the work as a feasibility study rather than a performance evaluation.", "body_md": "arXiv:2609.10790v1 Announce Type: cross \nAbstract: We present a Kubernetes Dynamic Resource Allocation (DRA) driver that makes composable CXL memory a schedulable cluster resource, and evaluate the resulting shared-memory tier for cross-node KV-cache reuse in LLM serving. The driver composes CXL regions on demand, materializes them as DAX devices on each participating host, and injects them into pods under a single Container Device Interface (CDI) name so that pods on different nodes access the same physical region. A shared-memory connector for vLLM/llm-d uses that region as a KV-cache tier with a slot directory embedded inside the shared medium, which eliminates the need for an external metadata service. On a two-node cluster with a 512\\,GiB CXL appliance and Qwen2.5-7B-Instruct, cross-node prefix reuse reduces TTFT by 5.5$\\times$--36.6$\\times$ at an external hit rate of 95.4--99.5\\,\\%, while node-local tiers (GPU prefix caching, CPU-DRAM offload) fall back to full recompute. The sharing gap, defined as the latency ratio between cross-node and same-node reuse, is 1--4\\%, indicating that cross-node reuse incurs little additional latency relative to same-node reuse on our testbed. Both replicas run full engines; the study demonstrates memory disaggregation rather than prefill/decode disaggregation. We report this as a feasibility study rather than a performance evaluation.", "url": "https://wpnews.pro/news/composable-cxl-memory-as-a-kubernetes-native-shared-memory-for-llm-serving", "canonical_source": "https://www.machinebrief.com/news/composable-cxl-memory-as-a-kubernetes-native-shared-memory-f-ijj3", "published_at": "2026-09-11 04:00:00+00:00", "updated_at": "2026-09-11 05:27:28.620663+00:00", "lang": "en", "topics": ["ai-infrastructure", "large-language-models", "ai-research", "mlops"], "entities": ["Kubernetes", "Compute Express Link (CXL)", "vLLM", "llm-d", "Qwen2.5-7B-Instruct", "arXiv"], "alternates": {"html": "https://wpnews.pro/news/composable-cxl-memory-as-a-kubernetes-native-shared-memory-for-llm-serving", "markdown": "https://wpnews.pro/news/composable-cxl-memory-as-a-kubernetes-native-shared-memory-for-llm-serving.md", "text": "https://wpnews.pro/news/composable-cxl-memory-as-a-kubernetes-native-shared-memory-for-llm-serving.txt", "jsonld": "https://wpnews.pro/news/composable-cxl-memory-as-a-kubernetes-native-shared-memory-for-llm-serving.jsonld"}}