cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 36/38 feed RSS

// recent coverage 748 mentions

00:00
2026-06-10
fergusfinn.com
large-language-models

Anatomy of a high-performance EP kernel

A high-performance Expert Parallelism (EP) kernel is essential for running large Mixture-of-Experts (MoE) language models across multiple GPUs, as it handles the dynamic routing of tokens to experts l…

12:29
2026-06-05
www2.eecs.berkeley.edu
large-language-models

vLLM: An Efficient Inference Engine for Large Language Models [pdf]

Researchers have released vLLM, a new inference engine designed to efficiently serve large language models by optimizing memory management and batching. The system achieves up to 24x higher throughput…

15:18
2026-06-04
github.com
large-language-models

KVarN: Native vLLM KV-cache quantization back end by Huawei

Huawei released KVarN, a native KV-cache quantization back end for vLLM that delivers up to 5x more cache capacity and 1.3x the throughput of FP16 while maintaining FP16-level accuracy. The calibratio…

17:27
2026-06-03
deeplearning.ai
large-language-models

Free vLLM Course: Inference, Compression, Benchmarks

DeepLearning.AI and Red Hat have released a free, intermediate-level course titled "Fast & Efficient LLM Inference with vLLM," taught by Red Hat Senior Developer Advocate Cedric Clyburn. The 1-hour 38…

00:00
2026-05-31
cefboud.com
large-language-models

Exploring Speculative Decoding: From Concept to Implementation

Speculative decoding optimizes LLM inference by using a cheap draft model to predict multiple tokens, which are then verified in a single forward pass of the target model, reducing memory-bandwidth bo…

20:38
2026-05-30
rajveerbachkaniwala.com
large-language-models

Stream2LLM: Overlap Context Streaming and Prefill for Reduced TTFT

Researchers have developed Stream2LLM, a system that extends the vLLM inference engine to support concurrent streaming of context to large language models, achieving up to 11x faster time-to-first-tok…

← prev page 36 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics