cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 26/38 feed RSS

// recent coverage 748 mentions

12:01
2026-07-22
pub.towardsai.net
large-language-models

The Complete Technical Guide to Running LLMs Locally in 2026

A technical guide to running large language models locally in 2026 provides hardware math, quantization tradeoffs, and benchmarks of five inference engines, with case studies from the author's 16GB Ap…

07:36
2026-07-22
leaddev.com
large-language-models

Your LLM inference benchmark is lying to you

Synthetic LLM inference benchmarks misrepresent production performance because they use fixed prompt lengths, steady request rates, and single-model hardware, while real traffic is bursty and variable…

03:48
2026-07-21
dev.to
artificial-intelligence

Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

Google's Gemma 4 E2B model serves efficiently on a single TPU v6e chip, achieving 213 tok/s for a single user and scaling to ~2,200 output tok/s across concurrent streams, while its QAT variants fail …

02:49
2026-07-21
arxiv.org
artificial-intelligence

Kimi Linear: An Expressive, Efficient Attention Architecture

Researchers at Moonshot AI introduced Kimi Linear, a hybrid linear attention architecture that outperforms full attention across short-context, long-context, and reinforcement learning scaling regimes…

07:08
2026-07-20
byteiota.com
artificial-intelligence

SQRL: Feyn’s Text-to-SQL Model Inspects Before It Writes

Feyn Labs shipped SQRL on July 19, a text-to-SQL model that inspects databases with read-only probes before writing queries, achieving 70.6% execution accuracy on BIRD Dev and edging past Claude Opus …

05:09
2026-07-20
byteiota.com
artificial-intelligence

Kimi K3 Open Weights Drop July 27: The Developer Prep Guide

Moonshot AI's Kimi K3, a 2.8-trillion-parameter model that topped the Frontend Code Arena leaderboard on day one, releases its full open weights on July 27. The MXFP4 weights require approximately 1.4…

10:42
2026-07-18
github.com
developer-tools

The Htop for LLM Inference

LLM Inspector, a new open-source CLI tool from developer Helal Saoudi, analyzes live LLM inference processes to show exactly how GPU memory is used by weights, KV cache, and workspace, then projects o…

18:07
2026-07-17
netflixtechblog.medium.com
large-language-models

In-House LLM Serving at Netflix

Netflix's AI Platform team built an in-house LLM serving stack, running the full pipeline from model deployment through inference inside its existing production environment. The team selected vLLM as …

← prev page 26 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics