cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 746

vLLM

mentions 746 type Organization page 10/38 feed RSS

// recent coverage 746 mentions

08:57
2026-09-04
forum.level1techs.com
artificial-intelligence

ROCm 10.0 on Polaris

A developer has ported ROCm 10.0 to AMD's gfx803 architecture (Polaris GPUs), enabling vLLM and llama.cpp to run on older cards like the RX 580. The port, which builds on earlier work with ROCm 6.4.4 …

21:34
2026-09-03
news.ycombinator.com
ai-agents

Show HN: Building AI agents client-side JavaScript

A developer has launched Buttercup, an open-source project demonstrating AI agents built entirely in client-side JavaScript that run in the browser, aiming to reduce infrastructure costs by eliminatin…

17:19
2026-09-03
dev.to
ai-infrastructure

Deploying Inference Using NVIDIA Dynamo and vLLM

NVIDIA Dynamo, an open-source inference framework, has been deployed with the vLLM backend to serve chat completion requests in both aggregated and disaggregated configurations. The deployment involve…

16:11
2026-09-03
promptcube3.com
artificial-intelligence

NVIDIA is pushing local inference speeds up by 1.

NVIDIA has introduced optimizations that boost local AI inference speeds by up to 1.9x, integrated directly into llama.cpp and vLLM, benefiting users of LM Studio and Ollama. The company also unveiled…

12:30
2026-09-03
wirt.ee
large-language-models

On-Prem LLM Inference

A technical guide details the production deployment of on-premises large language model inference using vLLM, covering an 8-GPU node running GLM-5.2/5.3 (NVFP4 MoE) and a single L40S running gemma-4-2…

12:01
2026-09-03
pub.towardsai.net
artificial-intelligence

Stop Wasting GPU Memory: A Deep Dive Into vLLM’s PagedAttention

VLLM's PagedAttention technique reduces GPU memory waste in LLM serving from 60-80% to less than 4% by partitioning the KV cache into non-contiguous blocks, according to a technical analysis. The meth…

00:46
2026-09-03
github.com
ai-infrastructure

Hot reload vLLM and sglang configs

A new open-source tool called trimtab lets operators change SGLang and vLLM scheduler settings live, cutting configuration changes from a 1-7 minute redeploy to about 15 ms, with zero dropped requests…

12:42
2026-09-01
miraflow.ai
large-language-models

Nemotron 3 Ultra Explained

On June 4, 2026, NVIDIA released Nemotron 3 Ultra, a 550 billion parameter open-weight hybrid Mamba-Attention Mixture-of-Experts model that activates about 55 billion parameters per token, designed fo…

12:55
2026-08-31
aptai.dev
ai-infrastructure

Show HN: A dedicated hub to find, test, and serve LLM adapters

AptAI launched a centralized hub for discovering, testing, and deploying LLM adapters, enabling one-click serverless deployment with sub-millisecond execution overhead and zero cold starts. The platfo…

09:00
2026-08-31
infoq.com
ai-infrastructure

Producing the World's Cheapest Tokens: A How-to Guide

Meryem Arik, co-founder of inference company Doubleword, said most companies overpay for AI inference by 2x to 5x, sometimes by an order of magnitude, because their inference stacks do not match their…

← prev page 10 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics