cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 4/23 feed RSS

// recent coverage 443 mentions

14:29
2026-08-12
empero.org
ai-agents

Abacus 0.6.0 — the agent that remembers being an agent

Abacus 0.6.0, the open-source coding agent from rethink, introduces persistent memory that improves over time, with six mechanisms including papercuts, memories, tethering, and a hive for delegation. …

13:59
2026-08-12
promptcube3.com
artificial-intelligence

Open weight AI is the only real hedge against a billionaire-led

Open-weight AI models are the only real hedge against a billionaire-led AI industry, according to a tech commentary, because they provide data sovereignty, latency control, and customization through f…

12:54
2026-08-12
promptcube3.com
large-language-models

Why my local LLM kept crashing during a RAG experiment

A developer's local LLM crashed during a RAG experiment due to CUDA out-of-memory errors caused by an oversized 32,768-token context window, which inflated the KV cache. Fixing the issue by sliding LM…

10:35
2026-08-11
twitter.com
ai-infrastructure

Inferact vLLM creators are hiring

The maintainers of the open-source vLLM inference engine at Inferact are hiring, according to a post praising the team as among the most skilled engineers globally. The post highlights their role in b…

09:13
2026-08-11
vincentschmalbach.com
large-language-models

What Is Batch Invariance in LLM Inference?

VLLM's batch-invariance documentation defines a feature ensuring a request produces the same inference result regardless of batch size, composition, request order, or scheduling under a fixed hardware…

00:00
2026-08-11
modelplane.ai
artificial-intelligence

Why Day 0 for Nemotron 3.5 Lightning wasn't a scramble

NVIDIA released Nemotron-3.5-Lightning, a 30B mixture-of-experts model with 3B active parameters, on the same day Modelplane, an open-source fleet-level control plane for inference, achieved zero-day …

00:00
2026-08-11
mindstudio.ai
artificial-intelligence

How to Run fuse-1 Lite Locally: VRAM, Setup, and Formats

Fuse-1 Lite, a 5.72B parameter mixture-of-experts coding model from LiquidAI, can run locally with VRAM needs ranging from 3.36 GB in 4-bit quantized form to about 12 GB in full bfloat16 precision, ac…

17:01
2026-08-10
promptcube3.com
artificial-intelligence

Why does Zuckerberg push for "imperfect" releases while other

Meta CEO Mark Zuckerberg advocates for releasing 'imperfect' AI models early to leverage real-world telemetry and community feedback, a strategy that turns users into a massive QA team and accelerates…

13:30
2026-08-10
cast.ai
artificial-intelligence

LLM Inference Cost Optimization: Run AI Inference for Less

Cast AI benchmark testing shows that continuous batching at batch size 8 reduces Llama 3.1 70B inference cost on a single H100 from approximately $0.60-$0.80 per million tokens to $0.15-$0.25 per mill…

13:08
2026-08-10
sourcefeed.dev
artificial-intelligence

Muse Glimmer Is Meta's Apology for the Llama License

Meta's Superintelligence Labs released Muse Glimmer, a 30-billion-parameter dense multimodal model with 131K context, under the Apache 2.0 license, marking a departure from the restrictive Llama Commu…

10:31
2026-08-10
twitter.com
artificial-intelligence

Muse Spark 1.2 (Meta) is reportedly becoming open-weight

Meta announced it will release an open-weight version of Muse Spark 1.2 and released Muse Glimmer, a 30B agentic model with open weights under Apache 2.0, which can run on 24GB of VRAM. The model is q…

← prev page 4 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics