cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 746

vLLM

mentions 746 type Organization page 9/38 feed RSS

// recent coverage 746 mentions

05:30
2026-09-09
helpnetsecurity.com
ai-safety

AI-Infra-Guard: Open-source security scanner for AI systems

Tencent's Zhuque Lab released AI-Infra-Guard, an open-source security scanner for AI systems that fingerprints services such as Ollama, vLLM and ComfyUI, checks them against more than 1,600 known CVEs…

22:09
2026-09-08
frontierroles.com
artificial-intelligence

Staff Applied AI Inference Engineer — Crusoe

Crusoe, a vertically integrated AI infrastructure company, is hiring a Staff Applied AI Inference Engineer in San Francisco with a salary range of $215,000–260,000 per year, which sits 8% above the $2…

09:00
2026-09-08
cohere.com
artificial-intelligence

Inside the megakernel serving engine for North Mini Code

Cohere released a serving engine for its North Mini Code model built around a decode megakernel that runs 1.25x to 1.41x faster than vLLM end-to-end on a single H100 with BF16 precision. The engine, a…

07:30
2026-09-08
snipvote.com
large-language-models

AMD covers 5 speculative decoding methods for vLLM on AMD GPUs

AMD and the vLLM project detailed five speculative decoding methods for vLLM on AMD GPUs, achieving up to 2.59x throughput improvement for certain large language models. The techniques enable producti…

00:00
2026-09-08
rocm.blogs.amd.com
artificial-intelligence

veRL on AMD: Production-Ready RL Post-Training on ROCm

AMD and the veRL project released a production-ready reinforcement learning post-training container for AMD Instinct GPUs, supporting MI300 and MI355 series on ROCm, with AITER-accelerated vLLM and SG…

13:01
2026-09-07
pub.towardsai.net
ai-infrastructure

Superlinked Inference Engine

Superlinked has released the Superlinked Inference Engine (SIE), an open-source server designed to consolidate the many small models used in AI agent workflows into a single deployment, addressing wha…

09:26
2026-09-07
vllm.ai
large-language-models

Speculative Decoding in vLLM on AMD GPUs

AMD's experiments with speculative decoding in vLLM on AMD Instinct MI300X and MI355X GPUs using the ROCm platform show that output-token throughput gains vary by drafting method, proposal length, mod…

05:00
2026-09-07
marktechpost.com
artificial-intelligence

IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B

The Institute of Foundation Models (IFM), the frontier lab launched by MBZUAI in May 2025, released K2 Horizon, a fleet of six Apache 2.0-licensed open-source models ranging from 0.9B to 375B paramete…

22:00
2026-09-04
systems.seas.harvard.edu
large-language-models

KV Cache on Flash: What to Write When Writes Wear Out

Anshvardhan Shetty, a first-year EECS undergraduate at Imperial College London and research intern with Professor Juncheng Yang at Harvard SEAS, presents a study of KV cache admission on high-bandwidt…

← prev page 9 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics