cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 5/23 feed RSS

// recent coverage 443 mentions

08:18
2026-08-10
ssenthilnathan3.github.io
machine-learning

MoE routing is just branch prediction

A software engineer's analysis argues that MoE routing in transformer inference is fundamentally the same problem as CPU branch prediction, and that KV cache management techniques such as prefix cachi…

00:11
2026-08-10
byteiota.com
developer-tools

Mirafold: Generative UI for Claude Code and Codex (2026)

Mirafold, an MIT-licensed open-source tool, launched a browser-based generative UI layer for terminal coding agents Claude Code, Codex, and Gemini CLI, turning raw agent output into live cards, depend…

19:17
2026-08-09
github.com
artificial-intelligence

LLM Scaler – LLM Support for Intel's Arc Pro B60 and B70 GPUs

Intel has released LLM Scaler, a GenAI solution for text, image, and video generation optimized for Intel Arc Pro B60 and B70 GPUs, with the latest version intel/llm-scaler-vllm:0.21.0-b2 adding Multi…

13:09
2026-08-09
sourcefeed.dev
artificial-intelligence

Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

An independent developer has released DeepSeek-V4-Flash-0731-Latent-Reasoning, a 284B-parameter MoE model with a 35.7M-parameter reasoning head that performs latent reasoning in a 1024-dimensional spa…

15:08
2026-08-08
sourcefeed.dev
artificial-intelligence

Your Agents Are Waiting on the CPU, Not the GPU

Red Hat published a post arguing that 'the CPU is back' for LLM inference, citing an Intel and Georgia Tech paper that found CPU-side tool processing accounts for 50–90% of total latency in agentic wo…

19:02
2026-08-05
tokenstead.ai
large-language-models

Pokee-Isaac 28B

Pokee AI released Pokee-Isaac 28B, a 28B-parameter proprietary non-decoder-only model claiming a 10M-token context that fits on a single RTX 4090 (24GB) in quantized form. Vendor-reported benchmarks i…

10:11
2026-08-05
byteiota.com
artificial-intelligence

Kimi K3 Open Weights: Self-Hosting Reality Check

Moonshot AI released the weights for Kimi K3 on July 27, a 2.8-trillion-parameter Mixture-of-Experts model with a 1-million-token context window, scoring third globally behind Claude Fable 5 Max and G…

02:49
2026-08-05
dev.to
large-language-models

Measuring LLM Prefix Caching: The Cache Hit Rate Metric

An engineer's benchmarking guide introduces a cache hit rate metric for measuring prefix caching effectiveness in LLM serving, implemented in the open-source tool llmperf-rs. The metric calculates the…

← prev page 5 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics