cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 442

vLLM

mentions 442 type Organization page 2/23 feed RSS

// recent coverage 442 mentions

21:26
2026-08-17
forum.level1techs.com
large-language-models

Thoughts from existing B70 users?

Existing B70 users report that the best Intel LLM performance is achieved using a specific GitHub gist, with Qwen3.8-27B reaching 2313.29 tokens/s prefill and 30.76 tokens/s generation on a B70. One u…

20:49
2026-08-17
ainexusdaily.vercel.app
large-language-models

How to Scale LLM Inference for AI Agents Using vLLM

FreeCodeCamp published a tutorial on scaling large language model inference for AI agents using vLLM, explaining GPU scheduling challenges and optimization techniques. The guide covers building intuit…

15:55
2026-08-17
techstrong.ai
artificial-intelligence

Red Hat Brings Enterprise AI at Scale Into Focus

Red Hat's head of product for AI platforms, Tushar Katarki, says open source models and platforms are improving faster than enterprises planned, giving organizations more control over cost, data, and …

00:00
2026-08-17
mindstudio.ai
artificial-intelligence

Qwen3.8-9B: Running the Community-Distilled Model Locally

Empero released Qwen3.8-9B, an unofficial community distillation of Alibaba's Qwen 3.8 model, compressing its reasoning into a 9B parameter model based on Qwen3.5-9B. Trained on roughly 70,000 reasoni…

20:09
2026-08-16
byteiota.com
artificial-intelligence

Meta Muse Glimmer: Run a 30B Coding Agent on Your GPU

Meta released Muse Glimmer on August 10, a 30B open-weight coding agent model that runs on consumer hardware, with quantized builds fitting in 24GB VRAM. The model, distilled from Meta's Muse Spark 1.…

01:57
2026-08-16
promptcube3.com
artificial-intelligence

Open source AI is shifting its center of gravity toward China

Open source AI development is increasingly centered in China, with models outperforming GPT-4 in coding and math benchmarks while being released under permissive licenses, according to a technical ana…

19:08
2026-08-15
byteiota.com
artificial-intelligence

Meta Muse Glimmer 30B: Local AI Agent on One GPU

Meta Superintelligence Labs released Muse Glimmer on August 10, a 30B-parameter open-weight model under Apache 2.0, designed for local agentic workflows and tool calling. The model, available on Huggi…

18:54
2026-08-15
promptcube3.com
large-language-models

Qwen 3.8 27B actually beats the larger 3.7 Plus in coding

Alibaba's Qwen 3.8 27B dense model outperforms the larger Qwen 3.7 Plus in coding and office productivity tasks, according to a technical review. The 27B model features a 262,000-token context window …

16:08
2026-08-15
sourcefeed.dev
artificial-intelligence

Netflix's LLM Ranker Just Beat Its Production System

Netflix's GenRec, an LLM-backed ranker in the 1B–10B parameter range, beat its mature production recommender system in a four-week A/B test on roughly 10% of live traffic, posting statistically signif…

10:15
2026-08-15
promptcube3.com
large-language-models

Alibaba's open source models just crossed 3 billion downloads

Alibaba's open-source Qwen 2.5 series has surpassed 3 billion downloads, according to the company. The models are praised for stability across quantization levels, coding proficiency, long-context han…

07:24
2026-08-15
promptcube3.com
artificial-intelligence

Since the provided content was only a title

Open-weight AI models have closed the performance gap with closed-source APIs, with specialized 7B to 30B models now outperforming older 175B models due to improved data quality and Mixture-of-Experts…

07:12
2026-08-15
byteiota.com
artificial-intelligence

Qwen3.8-27B Is Out: The Local AI Model Developers Need

Alibaba released Qwen3.8-27B, a 27.78-billion-parameter open-weights multimodal model under Apache 2.0, with a 262,144-token native context window and configurable reasoning. Vendor-reported benchmark…

← prev page 2 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics