cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 23/38 feed RSS

// recent coverage 748 mentions

15:08
2026-07-30
redhat.com
ai-infrastructure

Why self-hosted inference is essential

Red Hat AI warns that enterprises relying on third-party hosted APIs for AI agent inference undermine their own data sovereignty, as every prompt and tool call routes through external datacenters. The…

13:59
2026-07-30
openalternative.co
artificial-intelligence

LocalAI

LocalAI, a self-hosted runtime that runs AI workloads on user-controlled hardware, offers an OpenAI-compatible API supporting text generation, vision, speech, image and video generation, embeddings, a…

13:05
2026-07-30
github.com
artificial-intelligence

Kimi k3 run on RTX 5090

RunInfra enables running Kimi-Linear-48B, a distilled version of the full 2.78-trillion-parameter Kimi K3 model, on a single consumer GPU such as the RTX 5090 with 32 GB VRAM, achieving 113.83 tokens …

11:01
2026-07-30
promptcube3.com
large-language-models

Open-Weight Models Now Match Proprietary Titans

The accuracy gap between the best open-weight models and GPT-4o has shrunk to under 3% on structured data tasks, according to a developer's hands-on analysis. A fine-tuned Qwen2.5 72B model achieved 9…

22:08
2026-07-29
byteiota.com
artificial-intelligence

DeepSeek V4 Pro: 80.6% SWE-Bench at $0.87/M Output

DeepSeek V4 Pro, a Mixture-of-Experts model with 1.6 trillion total parameters, scores 80.6% on SWE-bench Verified, tying Gemini 3.1 Pro and achieving the highest score for any open-weight model. Pric…

12:00
2026-07-29
letsdatascience.com
artificial-intelligence

Microsoft Outlines a Three-Layer Agent Routing Stack for AKS

Microsoft's Azure Kubernetes Service engineering team published a reference implementation on June 29 that separates agent-request routing into three layers: semantic model selection via RouteLLM, gat…

07:06
2026-07-29
inmyhead.is
large-language-models

A Scheduler as a Lens into LLM Inference

A developer built a small scheduler in Go to understand vLLM's scheduler for LLM inference, tracing each design decision back to its vLLM equivalent. The scheduler operates in a tick loop with three p…

07:06
2026-07-29
inmyhead.is
artificial-intelligence

Preempting the Prefill, Part 3: Results & Benchmark

VLLM's slack-aware preemption policies rescued urgent request attainment at high load in a benchmark on 6× A100 SXM4 80GB GPUs running Llama 70B, where the control policy collapsed to 0% urgent attain…

05:46
2026-07-29
inmyhead.is
artificial-intelligence

Preempting the Prefill

A new paper, FlowPrefill by Hsieh et al., proposes preempting long LLM inference prefills mid-forward-pass to rescue urgent requests that would otherwise miss their time-to-first-token (TTFT) service-…

00:00
2026-07-29
cefboud.com
large-language-models

How Profitable is LLM Inference? Doing the Math on Kimi K3

LLM inference profitability depends on the trade-off between batch size and GPU count, which determines token latency and cost per million tokens. Applying this model to Kimi K3, which requires at lea…

00:00
2026-07-29
kondasamy.com
artificial-intelligence

Kimi K3: What a 2.8T Open Model Changes for Engineers

Moonshot AI released Kimi K3, a 2.78-trillion-parameter mixture-of-experts model with 104.2 billion active parameters, native vision, and a 1,048,576-token context window, claiming it is the first ope…

15:08
2026-07-28
sourcefeed.dev
artificial-intelligence

Linear Attention Just Graduated to Frontier Scale

Moonshot AI's Kimi Linear attention architecture, introduced in October, now powers the company's 2.8-trillion-parameter Kimi K3 flagship model released in mid-July, marking the first production deplo…

11:44
2026-07-28
modal.com
artificial-intelligence

What Is Flash Attention?

Flash Attention is an algorithm that speeds up training and inference of transformer models by using smart memory management on GPUs. The original version was released in 2022, followed by Flash Atten…

← prev page 23 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics