cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 449

vLLM

mentions 449 type Organization page 19/23 feed RSS

// recent coverage 449 mentions

15:00
2026-06-17
hiraditya.github.io
large-language-models

vLLM's op IR, or: where the inference engine meets the compiler

VLLM, a model-serving engine for large language models, introduced a small op-level IR to resolve the tension between acting as a compiler target and a hand-tuned kernel dispatcher. The IR allows vLLM…

06:33
2026-06-17
arxiv.org
machine-learning

Fearless Concurrency on the GPU

Researchers introduced cuTile Rust, a tile-based system for safe, idiomatic GPU kernel authoring in Rust that extends Rust's ownership discipline to GPU kernels. On the NVIDIA B200 GPU, cuTile Rust ac…

20:08
2026-06-16
byteiota.com
large-language-models

Mellum2: JetBrains Open-Sources a 12B MoE Coding Model

JetBrains open-sourced Mellum2, a 12B Mixture-of-Experts coding model under Apache 2.0, designed for air-gapped and compliance-locked environments where external API calls are prohibited. The model us…

19:37
2026-06-16
dev.to
large-language-models

Serving any LLM using a single command line with Flama

Flama 2.0 introduces first-class support for generative AI, enabling users to download, package, and serve large language models (LLMs) via a single command line. The framework allows fetching models …

18:11
2026-06-16
the-ai-corner.com
ai-infrastructure

Inference engineering is the 80% cost cut most teams miss

Inference engineering, the craft of optimizing GPU operations during AI model inference, can cut costs by up to 80% by addressing the split between prefill and decode phases. Two teams using the same …

17:32
2026-06-16
newsletter.semianalysis.com
machine-learning

RL Systems Mind the Gap: Matching Trainer and Generator Throughput

Anthropic CEO Dario Amodei said reinforcement learning shows the same log-linear scaling as pre-training, but RL system efficiency is critical to afford enough training. Experiments on open models sho…

08:45
2026-06-16
thecomputersciencebook.com
large-language-models

PagedAttention is more than virtual memory

PagedAttention, a memory optimization technique in the vLLM inference server, applies virtual memory concepts to manage the KV cache in large language models, improving throughput by reducing fragment…

05:21
2026-06-16
letsdatascience.com
large-language-models

CacheWise Improves KVCache Reuse for LLM Coding Agents

Researchers introduced CacheWise, a KVCache management layer for LLM coding agents, reducing evictions by 2-2.6x and improving session completion time by up to 3.5x in vLLM, according to a June 2026 a…

21:59
2026-06-15
github.com
ai-agents

Show HN: Phlox – Open-source self-hosted agentic web chat

Phlox, an open-source self-hosted agentic web chat application, has been released on GitHub. It supports any model provider including AWS Bedrock and OpenAI-compatible endpoints, and features agentic …

← prev page 19 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics