cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 743

vLLM

mentions 743 type Organization page 6/38 feed RSS

// recent coverage 743 mentions

19:13
2026-09-16
github.com
ai-agents

Neurogrid Terminal Coding Agent

NeuroGrid released nrgrd, a terminal user interface coding agent that connects to any OpenAI-compatible inference endpoint, including NeuroGrid Marketplace deployments, vLLM, Ollama, and LM Studio. Th…

19:07
2026-09-16
byteiota.com
artificial-intelligence

DeepSeek V4.1-Flash: MIT Weights and the Agent Cost Story

DeepSeek released V4.1-Flash on September 10 under an MIT license with weights on HuggingFace, pricing cache hits at $0.003 per million tokens off-peak versus $0.022 per million for V4-Pro. The Mixtur…

17:56
2026-09-16
promptcube3.com
ai-infrastructure

Vera Rubin NVL72 is hitting 3.7x the throughput of GB300 NVL72

NVIDIA's Vera Rubin NVL72 delivered up to 3.7x higher throughput than the GB300 NVL72 on the Qwen3-VL model in MLPerf Inference v6.1 preview submissions, using vLLM with the NVIDIA Dynamo framework, w…

20:55
2026-09-15
superml.dev
ai-infrastructure

Prefill, Not Decode, Is Your Agent's Real Bottleneck

Prefill, not decode, has become the dominant bottleneck in RAG and multi-agent workloads, according to an analysis citing NVIDIA's published figures of roughly 30x higher served-request counts for lar…

17:32
2026-09-15
twitter.com
ai-infrastructure

Show HN: Self-adjusting vLLM at production scale

Rivvr launched an autopilot product that automatically tunes vLLM inference deployments, claiming up to 2x higher tokens per second and 40-70% cuts in AWS bills. The company said the tool load-tests a…

16:00
2026-09-15
gladlabs.io
ai-infrastructure

Llama.cpp vs vLLM vs SGLang

Glad Labs decided not to switch its self-hosted inference stack from Ollama to vLLM after reviewing its own call logs, which showed only one to three concurrent calls at most against roughly 50 calls …

10:14
2026-09-14
dev.to
ai-tools

llama.cpp vs Ollama in 2026: Which Runtime Should You Run?

A technical comparison examines the tradeoffs between Ollama and llama.cpp for local LLM inference, framing the choice as one between a managed model service and a toolkit operated directly. The guide…

00:00
2026-09-14
mindstudio.ai
generative-ai

How to Run OUI-1 with vLLM for Generative UI

Thesys published OUI-1, a 26B-parameter diffusion model with 4B active parameters finetuned via LoRA from Google's DiffusionGemma 26B-A4B-it, which generates UI screens in openui-lang and scores 71.7%…

12:31
2026-09-13
pub.towardsai.net
large-language-models

Latency Optimization Levers for Open-Weight LLM inference:Part-2

A second-part technical article on latency optimization for open-weight LLM inference details four techniques measured on a fixed deployment of Qwen3-8B (Apache-2.0) served on vLLM through SageMaker's…

12:09
2026-09-13
sourcefeed.dev
ai-research

Transformers v5 turned a library into a standard

Hugging Face released Transformers v5.17.0 on September 9, adding seven model architectures in a single minor release, including Tencent's 780-billion-parameter mixture-of-experts model and Moonshot A…

← prev page 6 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics