cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 746

vLLM

mentions 746 type Organization page 12/38 feed RSS

// recent coverage 746 mentions

14:00
2026-08-28
kdnuggets.com
ai-infrastructure

The Local AI Stack for Productive SLMs

A practical guide from an unnamed author outlines a four-layer framework for building productive local AI stacks with small language models (1B–14B parameters), naming Ollama, LM Studio, llama.cpp, an…

06:02
2026-08-28
gist.github.com
large-language-models

Agent instructions for deploying Qwen 3.8 Flash Next

A developer documented a runbook for deploying the Qwen 3.8 Flash Next model on a single NVIDIA Blackwell GPU using a specialized vLLM runtime. The deployment supports both Docker and native systemd t…

02:30
2026-08-28
dev.to
large-language-models

The Best Self-Hosted LLMs in 2026 — and How I Deployed Them

A developer detailed how they deployed self-hosted open-weight LLMs for a fintech client whose compliance team banned external AI APIs, running a 70B-class model on two GPU boxes for document Q&A, sup…

00:00
2026-08-28
mindstudio.ai
artificial-intelligence

How to Run Tencent's Hy4 Preview Locally with vLLM or SGLang

Tencent's 770B-parameter Hy4 preview model, a Mixture-of-Experts architecture with 49B activated parameters per token, is now deployable locally via prebuilt Docker images for vLLM and SGLang, requiri…

18:01
2026-08-27
pub.towardsai.net
artificial-intelligence

Your Context Length Decides What a Kernel Is Worth

A 2× faster attention kernel yields only 0.66% end-to-end speedup on a 1,024-token prompt with a 128-token answer under vLLM's --goodput ttft:500 tpot:50 promise, according to a seven-part technical s…

16:18
2026-08-27
dev.to
ai-safety

The LLM Isn't Your Attacker. Your eval() Statement Is.

A critical vulnerability in vLLM's tool-call parser, CVE-2025-9141, allows model-generated arguments to be passed directly to eval(), enabling arbitrary code execution. The flaw highlights a broader p…

16:13
2026-08-27
servethehome.com
ai-infrastructure

Oxmiq Labs HBF in AI Compute at Hot Chips 2026

Oxmiq Labs presented high-bandwidth flash (HBF) as a capacity tier for AI inference at Hot Chips 2026, claiming it delivers 8 to 16 times the capacity of HBM at the same cost. The company detailed HBF…

17:09
2026-08-26
sourcefeed.dev
artificial-intelligence

Google's 4.7x Qwen 3.5 Speedup Is a Sharding Story

Google engineers reported a 4.7x faster prefill and 3.1x faster decode for Qwen 3.5-397B-A17B on Ironwood TPUs between April and June, achieved by using data parallelism for attention layers and exper…

← prev page 12 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics