cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 8/23 feed RSS

// recent coverage 443 mentions

12:00
2026-07-29
letsdatascience.com
artificial-intelligence

Microsoft Outlines a Three-Layer Agent Routing Stack for AKS

Microsoft's Azure Kubernetes Service engineering team published a reference implementation on June 29 that separates agent-request routing into three layers: semantic model selection via RouteLLM, gat…

07:06
2026-07-29
inmyhead.is
large-language-models

A Scheduler as a Lens into LLM Inference

A developer built a small scheduler in Go to understand vLLM's scheduler for LLM inference, tracing each design decision back to its vLLM equivalent. The scheduler operates in a tick loop with three p…

07:06
2026-07-29
inmyhead.is
artificial-intelligence

Preempting the Prefill, Part 3: Results & Benchmark

VLLM's slack-aware preemption policies rescued urgent request attainment at high load in a benchmark on 6× A100 SXM4 80GB GPUs running Llama 70B, where the control policy collapsed to 0% urgent attain…

05:46
2026-07-29
inmyhead.is
artificial-intelligence

Preempting the Prefill

A new paper, FlowPrefill by Hsieh et al., proposes preempting long LLM inference prefills mid-forward-pass to rescue urgent requests that would otherwise miss their time-to-first-token (TTFT) service-…

00:00
2026-07-29
kondasamy.com
artificial-intelligence

Kimi K3: What a 2.8T Open Model Changes for Engineers

Moonshot AI released Kimi K3, a 2.78-trillion-parameter mixture-of-experts model with 104.2 billion active parameters, native vision, and a 1,048,576-token context window, claiming it is the first ope…

00:00
2026-07-29
cefboud.com
large-language-models

How Profitable is LLM Inference? Doing the Math on Kimi K3

LLM inference profitability depends on the trade-off between batch size and GPU count, which determines token latency and cost per million tokens. Applying this model to Kimi K3, which requires at lea…

15:08
2026-07-28
sourcefeed.dev
artificial-intelligence

Linear Attention Just Graduated to Frontier Scale

Moonshot AI's Kimi Linear attention architecture, introduced in October, now powers the company's 2.8-trillion-parameter Kimi K3 flagship model released in mid-July, marking the first production deplo…

11:44
2026-07-28
modal.com
artificial-intelligence

What Is Flash Attention?

Flash Attention is an algorithm that speeds up training and inference of transformer models by using smart memory management on GPUs. The original version was released in 2022, followed by Flash Atten…

09:47
2026-07-28
comfyfile.com
large-language-models

How to Download and Run Kimi K3 Open Weights

Moonshot AI released the full Kimi K3 open weights on July 27, 2026, a 2.8-trillion-parameter mixture-of-experts model with a one-million-token context window and a 1.4 TB download. The model uses nat…

19:13
2026-07-27
promptcube3.com
large-language-models

Kimi K3: Why Open-Weight Models are Shaking Up the Market

Open-weight AI models like Kimi K3 are disrupting the market by enabling local deployment and customization without proprietary fees, according to the article. The shift lowers barriers for developers…

17:23
2026-07-27
cryptobriefing.com
artificial-intelligence

Moonshot completes Kimi K3 rollout with full model weight release

Moonshot AI released the full Kimi K3 model weights and technical report on July 27, 2026, making its 2.8 trillion parameter mixture-of-experts model available to developers and researchers. The model…

15:44
2026-07-27
vllm.ai
artificial-intelligence

Kimi K3 on vLLM: Up to 370 Tokens/sec

VLLM announces efficient day-0 support for Moonshot AI's Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model, achieving up to 370 tokens per second with speculative decoding on 16 NVIDIA GB300 …

15:16
2026-07-27
baseten.co
artificial-intelligence

We built a day-0 API for Kimi K3

Baseten has launched day-0 API support for Kimi K3, a new open frontier model from Moonshot AI with 2.8 trillion parameters, making it the largest open model to date. The API runs on NVIDIA GB300 NVL7…

← prev page 8 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics