cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 11/23 feed RSS

// recent coverage 443 mentions

03:48
2026-07-21
dev.to
artificial-intelligence

Gemma 4 E2B on a Single TPU v6e Chip: A Serving Deep Dive

Google's Gemma 4 E2B model serves efficiently on a single TPU v6e chip, achieving 213 tok/s for a single user and scaling to ~2,200 output tok/s across concurrent streams, while its QAT variants fail …

02:49
2026-07-21
arxiv.org
artificial-intelligence

Kimi Linear: An Expressive, Efficient Attention Architecture

Researchers at Moonshot AI introduced Kimi Linear, a hybrid linear attention architecture that outperforms full attention across short-context, long-context, and reinforcement learning scaling regimes…

07:08
2026-07-20
byteiota.com
artificial-intelligence

SQRL: Feyn’s Text-to-SQL Model Inspects Before It Writes

Feyn Labs shipped SQRL on July 19, a text-to-SQL model that inspects databases with read-only probes before writing queries, achieving 70.6% execution accuracy on BIRD Dev and edging past Claude Opus …

05:09
2026-07-20
byteiota.com
artificial-intelligence

Kimi K3 Open Weights Drop July 27: The Developer Prep Guide

Moonshot AI's Kimi K3, a 2.8-trillion-parameter model that topped the Frontend Code Arena leaderboard on day one, releases its full open weights on July 27. The MXFP4 weights require approximately 1.4…

18:00
2026-07-18
gist.github.com
artificial-intelligence

tutorial-kimi-k3.md

Moonshot AI's Kimi K3 model matches Opus 4.8 on intelligence benchmarks while costing ~70% less, with a 1M-token context window and open weights. The API is OpenAI-compatible, enabling easy integratio…

10:42
2026-07-18
github.com
developer-tools

The Htop for LLM Inference

LLM Inspector, a new open-source CLI tool from developer Helal Saoudi, analyzes live LLM inference processes to show exactly how GPU memory is used by weights, KV cache, and workspace, then projects o…

18:07
2026-07-17
netflixtechblog.medium.com
large-language-models

In-House LLM Serving at Netflix

Netflix's AI Platform team built an in-house LLM serving stack, running the full pipeline from model deployment through inference inside its existing production environment. The team selected vLLM as …

00:51
2026-07-17
tinfoil.sh
ai-infrastructure

Architecting Secure Prompt Caching

Tinfoil announces cached prompt pricing in its Inference API, a feature that reduces compute for eligible requests by caching recently processed inputs, but the company warns that caching introduces t…

16:11
2026-07-16
byteiota.com
artificial-intelligence

vLLM v0.25: Model Runner V2 Default, PagedAttention Gone

VLLM v0.25.0, released July 11, deletes the original PagedAttention implementation and makes Model Runner V2 the default execution backend for all dense models, delivering a 56% throughput improvement…

14:03
2026-07-16
github.com
artificial-intelligence

Veta: AI agent that QA-tests Android apps

Veta, an AI agent that QA-tests Android apps using a swarm of autonomous sub-agents, runs 100% of its AI inference on AMD GPUs via Fireworks AI on AMD Instinct or self-hosted vLLM on ROCm. The system …

← prev page 11 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics