cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 14/38 feed RSS

// recent coverage 748 mentions

05:40
2026-08-25
llmpanel.io
artificial-intelligence

LLMPanel Deploy vLLM to RunPod or Vast.ai Without Kubernetes

LLMPanel has launched an open-source platform that deploys large language models on any GPU cloud, including RunPod and Vast.ai, without requiring Kubernetes. The tool provisions containers, exposes O…

10:33
2026-08-24
dev.to
generative-ai

New advancements in Generative AI

A developer highlights the shift in generative AI tooling toward agentic workflows, local execution, and structured outputs. The post demonstrates how native structured outputs with JSON schemas and l…

07:00
2026-08-24
hiraditya.github.io
artificial-intelligence

A Bug Is a Violation of a Specification

A bug is a violation of a specification, and no specification exists that prefix caching's variable logits violate, according to an analysis of vLLM and SGLang issues. The vLLM PR #34046 adds an opt-i…

00:00
2026-08-24
rocm.blogs.amd.com
artificial-intelligence

Serving 64Mi-Token Contexts on One AMD Instinct™ MI355X Node

On a single 8-GPU AMD Instinct MI355X node, AMD served Kimi Linear 48B-A3B with context lengths from 1024 tokens to 64Mi (67,108,864), using vLLM with FP8 KV cache and tensor parallelism 8, recording …

22:51
2026-08-23
chipsandcheese.com
artificial-intelligence

Hot Chips 2026: Applying High Bandwidth Flash (HBF)

At Hot Chips 2026, Anurag Agarwal and Radhakrishna Giduthuri presented simulations and projections for High Bandwidth Flash (HBF), a new memory technology that combines SSD-like flash storage with HBM…

18:01
2026-08-23
pub.towardsai.net
large-language-models

Tuning vLLM: What Every Setting Does to the Arithmetic

VLLM v0.27.1, the open-source inference engine, has unified its scheduler around a fixed token budget per step, making chunked prefill, prefix caching, and speculative decoding stackable by default, a…

14:06
2026-08-23
weightless.msuiche.com
artificial-intelligence

Abliteration Without the Weights

A new 478 KB vector file enables 'abliteration without the weights' by removing refusal directions in activation space at inference time, allowing cybersecurity defenders to run capable models on thei…

07:21
2026-08-23
promptcube3.com
artificial-intelligence

Developers are losing their minds over a new open-source model

Developers are buzzing over a new open-source AI model distributed through decentralized channels with no official landing page or marketing campaign, according to a developer blog post. The anonymous…

06:26
2026-08-23
gojiberries.io
large-language-models

Cache-Control for LLMs

Anthropic's Claude Sonnet 5 charges $2.00 per million fresh input tokens, $2.50 to write into a five-minute cache, and $0.20 per read, making two identical calls cost $2.70 with caching versus $4.00 w…

05:36
2026-08-23
sankalp.bearblog.dev
large-language-models

How Prompt Caching Works

Prompt caching in large language model inference engines such as vLLM works by reusing key-value (KV) caches through paged attention and automatic prefix caching, enabling faster responses and lower c…

02:33
2026-08-23
github.com
developer-tools

Show HN: LayoutLens: AI-Powered Visual UI Testing

LayoutLens, an AI-powered visual UI testing tool, catches layout and accessibility bugs using deterministic axe-core and geometry checks that run keyless and free in CI, with an optional vision-LLM ti…

02:30
2026-08-23
dev.to
artificial-intelligence

Self-Hosted RAG: A Production Pipeline on Your Own Hardware

A developer detailed the construction of a fully self-hosted RAG pipeline for a Gulf bank that required no data to leave its premises. The project, built on two used 3090 GPUs and open-source tools li…

← prev page 14 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics