cd/entity/llm-d· home› entities› llm-d
grep -l @llm-d /news/*.json | wc -l → 14

llm-d

mentions 14 type Organization feed RSS

// recent coverage 14 mentions

15:01
2026-09-11
pub.towardsai.net
ai-infrastructure

LLM-D Explained

Llm-d is a Kubernetes-native distributed inference serving stack, released open source under Apache 2.0, that adds inference-aware orchestration on top of vLLM and Kubernetes without replacing either.…

22:32
2026-09-10
skeptrune.com
ai-infrastructure

The Inference Engineering Skills Map

A former B2B SaaS web developer who switched to inference engineering about a month ago published a skills map arguing the field requires only three service categories: routing and scheduling via Nvid…

12:00
2026-09-08
research.ibm.com
artificial-intelligence

How llm-d makes the most of the hardware you already have

IBM Research and Red Hat used the open-source llm-d framework to deploy GLM-5.2, an approximately 753-billion-parameter mixture-of-experts model, on 544 NVIDIA H100 GPUs, serving up to 3,000 concurren…

14:30
2026-08-19
hiraditya.github.io
artificial-intelligence

Two Schedulers, One SLO

A vLLM RFC from the llm-d team warns that disaggregated inference deployments, where prefill and decode run on separate schedulers, can trigger recomputation-based preemption inside the decode instanc…

15:55
2026-08-17
techstrong.ai
artificial-intelligence

Red Hat Brings Enterprise AI at Scale Into Focus

Red Hat's head of product for AI platforms, Tushar Katarki, says open source models and platforms are improving faster than enterprises planned, giving organizations more control over cost, data, and …

12:11
2026-08-11
byteiota.com
artificial-intelligence

llm-d Joins CNCF: Kubernetes Gets Serious About AI Inference

The Cloud Native Computing Foundation (CNCF) accepted llm-d, an open-source framework for distributed LLM inference on Kubernetes, as a sandbox project in March 2026, with contributions from IBM Resea…

16:00
2026-07-31
cloud.google.com
ai-infrastructure

What’s new in AI infrastructure and orchestration this month

Google Cloud announced the general availability of Managed Lustre, a high-performance storage solution powered by DDN's EXAScaler, offering throughput from 125 MB/s to 1000 MB/s per TiB and scaling up…

15:27
2026-06-27
cefboud.com
large-language-models

Distributed LLM Inference with LLM-d

A new open-source tool called llm-d acts as an LLM-aware load balancer for distributed inference, intelligently routing requests across vLLM instances based on KV cache locality and GPU utilization. B…

09:12
2026-06-24
byteiota.com
ai-infrastructure

NVIDIA Grove: Open-Source Kubernetes API for AI Inference

NVIDIA open-sourced Grove, a Kubernetes API for managing multi-component AI inference stacks, at KubeCon Europe 2026. Grove introduces custom resources for gang scheduling, topology-aware placement, a…

18:00
2026-06-23
research.ibm.com
artificial-intelligence

Running AI on mixed hardware for speed and affordability

IBM Research, Red Hat, and NxtGen Cloud Technologies demonstrated that using llm-d to serve AI models on mixed GPU hardware can boost inference speeds by 3 to 5 times and double throughput, enabling e…

// co-occurs with top 8 entities
// topics top 6 topics