cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 745

vLLM

mentions 745 type Organization page 8/38 feed RSS

// recent coverage 745 mentions

20:30
2026-09-10
neural-nova.com
large-language-models

Neural Nova – GPU optimization benchmarks for LLM workloads

Neural Nova published GPU optimization benchmarks for LLM workloads, reporting throughput and cost gains across four models running on vLLM. Qwen3-235B-A22B on 8× NVIDIA H100-80GB posted +138.7% token…

18:54
2026-09-10
frontierroles.com
ai-infrastructure

Inference Engineering, Co-op — Inferact

Inferact, founded by the creators and core maintainers of vLLM, is hiring University of Waterloo co-op students for an on-site Inference Engineering co-op in San Francisco, with pay not published. The…

13:39
2026-09-10
github.com
ai-infrastructure

LRU is harder to beat than the KV-cache papers suggest

A prefix-cache simulator replaying 68,266 requests from 393 real Claude Code sessions and 23,608 Mooncake requests failed to beat the production LRU baseline in three separate attempts, according to t…

18:09
2026-09-09
sourcefeed.dev
machine-learning

Tunix takes aim at the idle-TPU tax in agentic RL

Google's open-source, JAX-native post-training library Tunix now ships an agentic RL trainer built around asynchronous rollouts and a decoupled producer-consumer pipeline, aiming to eliminate idle TPU…

16:09
2026-09-09
sourcefeed.dev
artificial-intelligence

Ray's new TPU support is aimed at your GPU bill

Ray 2.55 adds official TPU support with atomic gang scheduling for TPU slices, aiming to cut GPU costs by enabling Ray-based workloads to run on Google's TPUs. Google's motive is to boost external TPU…

12:01
2026-09-09
pub.towardsai.net
large-language-models

Why Is Your LLM Recomputing the Same Prompt 1,000 Times a Day?

Production LLM traffic is dominated by repeated prompt prefixes—system prompts, chat history, and shared documents—causing inference engines to recompute identical KV cache state thousands of times da…

11:22
2026-09-09
blog.sentry.security
ai-safety

Living off Someone Else's Inference

Sentry's AI security research team, led by Armend Gashi and Redon Gashi, presented at DEF CON 34 a new tool called infreerence that demonstrates how adversaries exploit exposed self-hosted inference s…

05:56
2026-09-09
forum.level1techs.com
artificial-intelligence

Local LLM Based Coding on HPC Equipment

A developer running custom simulation software has moved from cloud-based AI coding assistance to local inference on high-end HPC equipment, using Hermes, Ollama, vLLM, and SGLang with models like Qwe…

05:30
2026-09-09
helpnetsecurity.com
ai-safety

AI-Infra-Guard: Open-source security scanner for AI systems

Tencent's Zhuque Lab released AI-Infra-Guard, an open-source security scanner for AI systems that fingerprints services such as Ollama, vLLM and ComfyUI, checks them against more than 1,600 known CVEs…

← prev page 8 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics