cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 35/38 feed RSS

// recent coverage 748 mentions

04:43
2026-06-14
github.com
large-language-models

Forked TensorZero after it was archived after raising $7.3M

Agentify has forked the archived TensorZero project, which raised $7.3M, and released Agentify Gateway, an open-source LLM gateway with observability, evaluation, optimization, and experimentation fea…

01:13
2026-06-14
byteiota.com
large-language-models

DiffusionGemma: Google’s 4x Faster Text Diffusion Model

Google DeepMind released DiffusionGemma on June 10, 2026, a 26B open-weight text diffusion model that generates 256 tokens simultaneously, achieving up to 1,008 tokens per second on an H100—4-5x faste…

17:51
2026-06-12
testingcatalog.com
artificial-intelligence

MiniMax M3 launches on NVIDIA platform with Free Endpoint

MiniMax released its M3 multimodal model on NVIDIA's accelerated infrastructure, offering a free public endpoint via NVIDIA's API catalog. The 428-billion-parameter model processes text, images, and v…

17:20
2026-06-11
developers.googleblog.com
large-language-models

DiffusionGemma: The Developer Guide

Google has released DiffusionGemma, an experimental text-generation model built on the Gemma 4 architecture that generates text in parallel blocks rather than token-by-token, enabling faster inference…

17:00
2026-06-10
pytorch.org
large-language-models

Portable vLLM Model Inference Kernels in Helion

Helion kernels were integrated into vLLM for FP8 inference using Qwen3 models and evaluated across NVIDIA H100 and B200 GPUs. The experiments demonstrated that Helion provides a productive PyTorch-nat…

16:25
2026-06-10
phoronix.com
artificial-intelligence

AMD's Lemonade SDK For Local AI Adds NVIDIA CUDA Support

AMD released Lemonade SDK version 10.7, adding NVIDIA CUDA support to its local AI server solution that previously only supported AMD hardware, Apple Metal GPUs, and AArch64 CPUs. The update integrate…

← prev page 35 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics