cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 743

vLLM

mentions 743 type Organization page 2/38 feed RSS

// recent coverage 743 mentions

11:44
2026-09-29
x.com
large-language-models

Qwen3.8-Flash-Next Is on TensorFold with Speed Boosts

TensorFold 0.3.6.2 delivered decode speeds over 62 tokens per second on a single stream and 119 tokens per second across five concurrent streams running Qwen3.8-Flash-Next on a single Nvidia DGX Spark…

03:13
2026-09-29
developers.redhat.com
large-language-models

Run Decision Models on vLLM and Red Hat AI Using DiffusionGemma

The vLLM community added a structured-read mode to DiffusionGemma 26B-A4B via vLLM PR #57250, letting the open model return typed probabilistic decisions in a single denoising step instead of token-by…

00:00
2026-09-28
int21.ai
ai-infrastructure

An AlphaGo Moment for Inference?

INT21 generated 20 inference engines across seven model categories in two weeks using Rust with C++ and CUDA components, and its MiMo engine reached 1,308 tokens/s versus 540 for tuned SGLang and 1,01…

00:00
2026-09-28
mindstudio.ai
large-language-models

OrcaSAQ2 27B: Run Qwen3.8-27B in 12GB With 3-Bit Quantization

OrcaRouter released OrcaSAQ2 27B, a 3-bit mixed-precision quantization of Qwen3.8-27B that shrinks the checkpoint from 54GB in BF16 to 12.3GB, a 77.2% reduction, under Apache-2.0. The model card repor…

10:41
2026-09-27
stackness.dev
ai-tools

How do you run System One decision models locally?

Ollaya reached release 0.7.3 on 27 September 2026, four days after its repository appeared on 23 September, and now serves ten model families behind an API it calls wire-identical to TypeSafe's Jev, a…

22:01
2026-09-26
pub.towardsai.net
ai-safety

[Framework] Zero Standing Privilege for AI Workloads

Enterprise data breaches involving unsanctioned Shadow AI cost organizations an average of $650,000 more than conventional security incidents, and 20.0% of global enterprises have already suffered a p…

19:57
2026-09-26
jadidbourbaki.github.io
large-language-models

42x faster prompt lookup drafting in llama.cpp

A developer writing as jadidbourbaki reports making prompt lookup decoding drafting in llama.cpp up to 42x faster while using up to 2.6x less memory, through performance optimizations based on work by…

← prev page 2 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics