cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 444

vLLM

mentions 444 type Organization page 17/23 feed RSS

// recent coverage 444 mentions

18:00
2026-06-23
research.ibm.com
artificial-intelligence

Running AI on mixed hardware for speed and affordability

IBM Research, Red Hat, and NxtGen Cloud Technologies demonstrated that using llm-d to serve AI models on mixed GPU hardware can boost inference speeds by 3 to 5 times and double throughput, enabling e…

00:00
2026-06-23
modelplane.ai
ai-infrastructure

Introducing Modelplane: the control plane for AI inference

Modelplane, an open-source control plane for AI inference built on Crossplane, is being released to manage GPU clusters as a single inference fleet, handling provisioning, model placement, autoscaling…

11:47
2026-06-21
vettedconsumer.com
large-language-models

Show HN: Local LLM Hardware Calculator

A new Local LLM Hardware Calculator helps users estimate memory requirements for running large language models on their own hardware, factoring in weights, KV cache, and overhead. The tool also compar…

14:24
2026-06-20
news.ycombinator.com
ai-infrastructure

How to become an AI infrastructure engineer?

An AI infrastructure engineer at a major industrial company seeks advice on transitioning from SRE-focused work to a proper software engineering role in AI infrastructure, asking for skills, resources…

01:36
2026-06-20
dev.to
large-language-models

KV cache and PagedAttention: what they do and why they matter

A developer explains that the KV cache is the biggest operational bottleneck in production LLM serving on GPUs, consuming more memory than model weights for workloads with high concurrency or long con…

00:17
2026-06-20
modal.com
large-language-models

Speculation Is All You Need

Modal Labs released state-of-the-art DFlash speculators for Qwen 3.5 and Qwen 3.6 models on Hugging Face, achieving 5-20% additional speedups and enabling Qwen 3.5 122B-A10B to run at over 1000 tok/s …

15:00
2026-06-19
hiraditya.github.io
large-language-models

Building vLLM from Source: A Field Guide (with all the pitfalls)

A developer building vLLM from source on an AWS g5 instance with Ubuntu 26.04 and Python 3.14 encountered multiple version-skew, driver, and toolchain issues, including a pitfall where missing nvidia-…

10:44
2026-06-19
discuss.huggingface.co
large-language-models

Gemma 4 bug fixes and Research Request

A critical bug in Google's Gemma 4 causes it to malform tool calls under real load, affecting vLLM, llama.cpp, Ollama, and oobabooga. A developer open-sourced a diagnosis, repair, and experimental LoR…

← prev page 17 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics