cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 20/38 feed RSS

// recent coverage 748 mentions

13:30
2026-08-10
cast.ai
artificial-intelligence

LLM Inference Cost Optimization: Run AI Inference for Less

Cast AI benchmark testing shows that continuous batching at batch size 8 reduces Llama 3.1 70B inference cost on a single H100 from approximately $0.60-$0.80 per million tokens to $0.15-$0.25 per mill…

13:08
2026-08-10
sourcefeed.dev
artificial-intelligence

Muse Glimmer Is Meta's Apology for the Llama License

Meta's Superintelligence Labs released Muse Glimmer, a 30-billion-parameter dense multimodal model with 131K context, under the Apache 2.0 license, marking a departure from the restrictive Llama Commu…

10:31
2026-08-10
twitter.com
artificial-intelligence

Muse Spark 1.2 (Meta) is reportedly becoming open-weight

Meta announced it will release an open-weight version of Muse Spark 1.2 and released Muse Glimmer, a 30B agentic model with open weights under Apache 2.0, which can run on 24GB of VRAM. The model is q…

08:18
2026-08-10
ssenthilnathan3.github.io
machine-learning

MoE routing is just branch prediction

A software engineer's analysis argues that MoE routing in transformer inference is fundamentally the same problem as CPU branch prediction, and that KV cache management techniques such as prefix cachi…

00:11
2026-08-10
byteiota.com
developer-tools

Mirafold: Generative UI for Claude Code and Codex (2026)

Mirafold, an MIT-licensed open-source tool, launched a browser-based generative UI layer for terminal coding agents Claude Code, Codex, and Gemini CLI, turning raw agent output into live cards, depend…

00:00
2026-08-10
sailresearch.com
ai-infrastructure

HTDYM (How To Deploy Your Model)

Sail Research has open-sourced HTDYM (How To Deploy Your Model), an internal performance modeling infrastructure that helps determine the optimal hardware and sharding configuration for serving large …

19:17
2026-08-09
github.com
artificial-intelligence

LLM Scaler – LLM Support for Intel's Arc Pro B60 and B70 GPUs

Intel has released LLM Scaler, a GenAI solution for text, image, and video generation optimized for Intel Arc Pro B60 and B70 GPUs, with the latest version intel/llm-scaler-vllm:0.21.0-b2 adding Multi…

13:09
2026-08-09
sourcefeed.dev
artificial-intelligence

Latent Reasoning Escapes the Lab, Bolted Onto DeepSeek-V4

An independent developer has released DeepSeek-V4-Flash-0731-Latent-Reasoning, a 284B-parameter MoE model with a 35.7M-parameter reasoning head that performs latent reasoning in a 1024-dimensional spa…

15:08
2026-08-08
sourcefeed.dev
artificial-intelligence

Your Agents Are Waiting on the CPU, Not the GPU

Red Hat published a post arguing that 'the CPU is back' for LLM inference, citing an Intel and Georgia Tech paper that found CPU-side tool processing accounts for 50–90% of total latency in agentic wo…

19:02
2026-08-05
tokenstead.ai
large-language-models

Pokee-Isaac 28B

Pokee AI released Pokee-Isaac 28B, a 28B-parameter proprietary non-decoder-only model claiming a 10M-token context that fits on a single RTX 4090 (24GB) in quantized form. Vendor-reported benchmarks i…

← prev page 20 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics