cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 441

vLLM

mentions 441 type Organization page 1/23 feed RSS

// recent coverage 441 mentions

15:54
2026-08-19
hetzner.com
artificial-intelligence

Hetzner Inference Experiment: Open-weights SLMs free of charge

Hetzner Online GmbH launched its Inference Experiment, a free inference API for open-weights small language models hosted on its infrastructure, one week ago. The company reported that small models li…

09:08
2026-08-19
sourcefeed.dev
developer-tools

Stop Shipping MCP Servers That Need a Python Runtime

Rust SDK rmcp 3.x now tracks the stable 2026-07-28 MCP spec revision and offers compile-time schema checking and single-binary distribution, making Rust a more reliable choice for MCP servers that hav…

00:20
2026-08-19
inco.ai
artificial-intelligence

DFlash 2: Keep Drafting Parallel

Inco AI released DFlash 2, a parallel speculative decoding technique that delivers over 20% more output from every verification pass with around 1% added cycle latency, achieving 2.7–3.4× throughput o…

23:21
2026-08-18
baseten.co
ai-infrastructure

Inference Engineering by Philip Kiely – Digital Download

Philip Kiely's new book, 'Inference Engineering,' is now available as a digital download, offering a comprehensive guide to the technologies and techniques powering AI inference across runtime, infras…

22:34
2026-08-18
promptcube3.com
large-language-models

Llama 3.

Meta's Llama 3.1 70B model can now run on a single 24GB consumer GPU using GGUF or EXL2 quantization, achieving 5-10 tokens per second on an RTX 3090, according to a deployment guide. The guide recomm…

14:30
2026-08-18
hiraditya.github.io
artificial-intelligence

The KV Cache Has No ABI

The KV cache has no standard ABI, with vLLM's FlashAttention backend alone reporting its cache shape as a four-dimensional tensor that varies by backend, attention variant, and model family, complicat…

page 1 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics