cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 747

vLLM

mentions 747 type Organization page 13/38 feed RSS

// recent coverage 747 mentions

20:13
2026-08-25
superml.dev
artificial-intelligence

Why Agentic Workloads Break Your Inference Stack

NVIDIA announced on August 24, 2026, that its Groq 3 LPX chip, built on technology licensed from Groq for a reported $20 billion, has entered full production as a dedicated decode extension to its Ver…

18:02
2026-08-25
dev.to
ai-agents

What Hermes Agent Gets Right About Long Running Agents

Nous Research's open source Hermes Agent runtime, released in February 2026 under the MIT license, is designed to close the gap in long-running agents by using a three-phase process that includes a re…

16:31
2026-08-25
pub.towardsai.net
large-language-models

The Ultimate Guide to LLM Inference Optimization- Part 2

The second part of a guide to LLM inference optimization focuses on attention mechanisms and KV cache management, explaining the compute-bound prefill phase and memory-bound decode phase. It introduce…

15:14
2026-08-25
huggingface.co
large-language-models

Granite 4.2 LLMs: How They're Built

IBM's Granite Team released Granite 4.2, a family of dense, decoder-only reasoning LLMs in 3B, 8B, and 30B sizes, pre-trained from scratch on roughly 15 trillion tokens with a five-phase strategy exte…

07:00
2026-08-25
hiraditya.github.io
artificial-intelligence

The Illusion of Determinism in Disaggregated Inference

SGLang and vLLM, the two major LLM serving engines, have attempted to enforce deterministic inference but face fundamental challenges due to floating-point non-associativity in GEMM kernels, which cau…

06:16
2026-08-25
byteiota.com
artificial-intelligence

Kimi K3: The Open-Weight Frontier Model Devs Should Know

Moonshot AI's Kimi K3, an open-weight model with 2.8 trillion total parameters but only 104 billion active per token, outperforms proprietary rivals on agentic coding benchmarks, scoring 42.0 on SWE M…

← prev page 13 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics