cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 743

vLLM

mentions 743 type Organization page 3/38 feed RSS

// recent coverage 743 mentions

15:49
2026-09-26
privatemode.ai
large-language-models

Turning GLM-5.3-Flash into a Jev-like decision model

Privatemode researchers Johannes Hötter and Marko Rosenmüller demonstrated that the off-the-shelf LLM GLM-5.3-Flash can make typed decisions with a probability for every option in a single forward pas…

00:00
2026-09-26
mindstudio.ai
artificial-intelligence

MiMo-V2.6-Flash-RL: Xiaomi's Efficient 309B Omnimodal Model

Xiaomi released MiMo-V2.6-Flash-RL, a sparse Mixture-of-Experts omnimodal model with 309 billion total parameters and 15 billion active per token, positioned as the efficiency-focused sibling to MiMo-…

00:00
2026-09-26
mindstudio.ai
artificial-intelligence

How to Run Audio8 ASR Infinite Locally with vLLM or Docker

Edge0 released Audio8 ASR Infinite, an open-weight streaming speech recognition model under the Apache 2.0 license, available on Hugging Face as Edge0/Audio8-ASR-Infinite and on GitHub as Edge0-AI/Aud…

00:00
2026-09-25
mindstudio.ai
artificial-intelligence

MiMo-V2.6-Pro-RL: Xiaomi's 1T-Parameter Agentic Model, Explained

Xiaomi's MiMo team released MiMo-V2.6-Pro-RL, an open-weight trillion-parameter mixture-of-experts agentic model with 1.02 trillion total parameters, 42 billion active per token, and a 1 million token…

09:20
2026-09-24
openalternative.co
large-language-models

vLLM

VLLM is an open-source large language model serving engine built around PagedAttention, which manages the KV cache the way an operating system manages virtual memory, and continuous batching, which ke…

12:48
2026-09-23
gist.github.com
ai-agents

How I build software with coding agents

A developer detailed a year-long setup for running coding agents like Claude Code and Codex CLI on real products, built around a single 32-core, 128 GB Linux box reached over Tailscale and managed wit…

10:21
2026-09-23
parity.io
ai-infrastructure

Self-hosting DeepSeek V4 for a software engineering org

Parity engineers ran a self-hosted inference trial of the open-weight DeepSeek V4 Flash model that handled more than 144,000 requests and almost 12.9 billion tokens from 25 engineers between 16 August…

08:22
2026-09-23
agentsearchengine.app
ai-agents

Qwen Code shipped an update

Qwen Code, the Qwen team's open-source terminal coding agent, shipped an update and now lists 28k GitHub stars, with its last update recorded in September 2026. The TypeScript project, licensed Apache…

← prev page 3 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics