cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 32/38 feed RSS

// recent coverage 748 mentions

10:01
2026-06-25
discuss.huggingface.co
large-language-models

Deepseek? Qwen?

A single H200 GPU with 141GB HBM3e cannot comfortably run DeepSeek V4 Flash (284B total, 13B active parameters) due to VRAM constraints, even with 2TB system RAM for offloading. The model requires an …

09:50
2026-06-25
oracomputing.com
large-language-models

ORA: Smaller Models. Same Intelligence

Ora Computing launched an automated LLM compression engine that reduces model size by up to 70% with minimal accuracy loss, enabling deployment on edge devices, on-prem servers, or cloud infrastructur…

20:51
2026-06-24
blog.crossplane.io
ai-infrastructure

I built a fleet-scale inference control plane using Crossplane

A developer built Modelplane, an open-source inference control plane using Crossplane, to manage GPU fleets across clouds, neoclouds, and on-premise environments. The platform allows platform teams to…

18:00
2026-06-23
research.ibm.com
artificial-intelligence

Running AI on mixed hardware for speed and affordability

IBM Research, Red Hat, and NxtGen Cloud Technologies demonstrated that using llm-d to serve AI models on mixed GPU hardware can boost inference speeds by 3 to 5 times and double throughput, enabling e…

00:00
2026-06-23
modelplane.ai
ai-infrastructure

Introducing Modelplane: the control plane for AI inference

Modelplane, an open-source control plane for AI inference built on Crossplane, is being released to manage GPU clusters as a single inference fleet, handling provisioning, model placement, autoscaling…

11:47
2026-06-21
vettedconsumer.com
large-language-models

Show HN: Local LLM Hardware Calculator

A new Local LLM Hardware Calculator helps users estimate memory requirements for running large language models on their own hardware, factoring in weights, KV cache, and overhead. The tool also compar…

14:24
2026-06-20
news.ycombinator.com
ai-infrastructure

How to become an AI infrastructure engineer?

An AI infrastructure engineer at a major industrial company seeks advice on transitioning from SRE-focused work to a proper software engineering role in AI infrastructure, asking for skills, resources…

01:36
2026-06-20
dev.to
large-language-models

KV cache and PagedAttention: what they do and why they matter

A developer explains that the KV cache is the biggest operational bottleneck in production LLM serving on GPUs, consuming more memory than model weights for workloads with high concurrency or long con…

← prev page 32 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics