cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 746

vLLM

mentions 746 type Organization page 11/38 feed RSS

// recent coverage 746 mentions

09:00
2026-08-31
infoq.com
ai-infrastructure

Producing the World's Cheapest Tokens: A How-to Guide

Meryem Arik, co-founder of inference company Doubleword, said most companies overpay for AI inference by 2x to 5x, sometimes by an order of magnitude, because their inference stacks do not match their…

00:00
2026-08-31
neuronpedia.org
ai-research

Introducing: interp-engine 🚀🔎

Decode Research has open-sourced interp-engine, a high-performance interpretability engine built from scratch to run production Neuronpedia workloads, including Jacobian Lens, NLAs, circuit tracing, a…

19:39
2026-08-30
dev.to
artificial-intelligence

Tencent Releases WeMM-Embedding for Multimodal Retrieval

Tencent's WeChat Vision team released WeMM-Embedding, a family of multimodal embedding models in 2B, 4B, and 9B parameter sizes, supporting text, images, videos, visual documents, and interleaved inpu…

17:47
2026-08-30
dev.to
ai-safety

AI Model Security Training: What a Platform Must Teach

A developer warns that AI model security training often overlooks the risk of malicious PyTorch checkpoints, which are programs executed via pickle deserialization, citing CVE-2025-24357 and CVE-2024-…

06:49
2026-08-30
dev.to
ai-safety

The Hidden Security Blind Spots in Local AI Workflows

A developer has identified critical security blind spots in local AI workflows, including unauthenticated network exposure when tools like Ollama and vLLM bind to 0.0.0.0, clipboard API key leakage, a…

00:07
2026-08-30
dev.to
ai-infrastructure

Standing Up a GPU Cluster on AKS for vLLM

Josef Doornink, an engineer, published a guide to standing up a GPU cluster on Azure Kubernetes Service (AKS) for serving vLLM models. The walkthrough covers requesting GPU quota, creating a cluster w…

20:21
2026-08-29
dev.to
artificial-intelligence

The State of Conversational AI in 2026, measured

A developer-led project has released a vendor-neutral report on the state of conversational AI in 2026, built entirely from primary data sources such as GitHub, Hugging Face, and search demand. Key fi…

11:11
2026-08-29
byteiota.com
ai-infrastructure

Samsung LPDDR5X-PIM at Hot Chips 2026: Developer Guide

Samsung unveiled LPDDR5X-PIM at Hot Chips 2026, a drop-in DRAM replacement that embeds multiply-accumulate units in each of its 16 memory banks, achieving 81.3 tokens per second on Llama 3.1 8B versus…

00:00
2026-08-29
digitalapplied.com
artificial-intelligence

When Your AI Agent Quietly Gets a Half-Finished Answer

A new technical reference documents that AI model APIs, including Anthropic, OpenAI, and vLLM, return HTTP 200 with well-formed bodies when generation hits the max output token ceiling, making truncat…

21:30
2026-08-28
pytorch.org
artificial-intelligence

vLLM Sessions at PyTorch Conference North America 2026

PyTorch Conference North America 2026, held October 20–21 in San Jose, CA, will feature vLLM across multiple sessions on KV cache management, disaggregated serving, hardware portability, kernel optimi…

20:50
2026-08-28
forum.level1techs.com
large-language-models

HP Z8 Fury G6i -- Perfect Qwen Flash Next setup (2x RTX Pro 6000s)

HP's Z8 Fury G6i workstation, equipped with dual RTX PRO 6000 Blackwell GPUs (96 GB each), 125 GB RAM, and a 48-core Intel Xeon X658X, serves Qwen3.8-Flash-Next-FP8 (125B MoE) at up to 218 tokens/s si…

20:01
2026-08-28
a11ce.com
large-language-models

Running Llama 3.1 405B

Meta's Llama 3.1 405B model is no longer hosted by any public inference provider, according to a guide on a11ce.com, which provides instructions for running the model on an on-demand GPU instance for …

← prev page 11 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics