cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 743

vLLM

mentions 743 type Organization page 7/38 feed RSS

// recent coverage 743 mentions

11:42
2026-09-12
dev.to
large-language-models

ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide

A 2026 guide compares ROCm and Vulkan as backends for hosting local LLMs on AMD GPUs, concluding that the choice depends on the inference engine, GPU generation, and workload rather than being interch…

04:11
2026-09-12
skydiscover-ai.github.io
ai-agents

Building Specialized Systems We Can Trust with Agents

SkyDiscover-Synthesize (SkySynth), an autonomous pipeline that builds just-in-time specialized systems, synthesizes key-value stores up to 2.3× faster than Redis and FASTER, inference engines with up …

20:09
2026-09-11
forum.level1techs.com
large-language-models

LLM Development System (GPUs Accounted for)

A user testing a multi-GPU LLM development machine reported that the system held up well under sustained load running the Qwen3.8 27B model, despite the machine shipping from 45 Drives with the wrong …

18:06
2026-09-11
dev.to
large-language-models

Why LLM Load Tests Are Costing You Thousands

A developer detailed how load testing LLM API integrations can cost thousands of dollars in metered tokens, citing a $3,000 bill from a failed 100,000-request test. The engineer, who built an Autonomo…

15:01
2026-09-11
pub.towardsai.net
ai-infrastructure

LLM-D Explained

Llm-d is a Kubernetes-native distributed inference serving stack, released open source under Apache 2.0, that adds inference-aware orchestration on top of vLLM and Kubernetes without replacing either.…

09:09
2026-09-11
dev.to
large-language-models

KV Cache on 16 GB GPUs: Making Long Context Actually Fit

A developer published a practical guide to fitting long-context LLM inference on 16 GB GPUs by budgeting VRAM for the KV cache, which grows with every active token and sequence. The guide provides a K…

06:15
2026-09-11
kagifeedback.org
ai-policy

Does Kagi watermark AI outputs (or plan to)?

A Kagi user publicly asked the privacy-focused search company to clarify whether outputs from its AI features — Assistant, Universal Summarizer, FastGPT, and Translate — are watermarked or otherwise m…

00:00
2026-09-11
mindstudio.ai
ai-infrastructure

RunPod Serverless Explained: Pay-Per-Second GPU API Deployment

RunPod Serverless lets developers deploy any Hugging Face model as an OpenAI-compatible API endpoint with per-second billing, scale-to-zero, and flash boot cold starts, according to RunPod. The platfo…

23:44
2026-09-10
forum.level1techs.com
ai-infrastructure

Post / Show off your Ai Rig - Whatchya got in there?

A user detailed a home AI inference rig built on an AMD Epyc Rome 7532 CPU with 128GB DDR4, roughly 5TB of NVMe storage, and 4x AMD V620 GPUs on an open-air crypto-mining bench. The rig runs a LiteLLM…

22:32
2026-09-10
skeptrune.com
ai-infrastructure

The Inference Engineering Skills Map

A former B2B SaaS web developer who switched to inference engineering about a month ago published a skills map arguing the field requires only three service categories: routing and scheduling via Nvid…

← prev page 7 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics