cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 33/38 feed RSS

// recent coverage 748 mentions

00:17
2026-06-20
modal.com
large-language-models

Speculation Is All You Need

Modal Labs released state-of-the-art DFlash speculators for Qwen 3.5 and Qwen 3.6 models on Hugging Face, achieving 5-20% additional speedups and enabling Qwen 3.5 122B-A10B to run at over 1000 tok/s …

15:00
2026-06-19
hiraditya.github.io
large-language-models

Building vLLM from Source: A Field Guide (with all the pitfalls)

A developer building vLLM from source on an AWS g5 instance with Ubuntu 26.04 and Python 3.14 encountered multiple version-skew, driver, and toolchain issues, including a pitfall where missing nvidia-…

10:44
2026-06-19
discuss.huggingface.co
large-language-models

Gemma 4 bug fixes and Research Request

A critical bug in Google's Gemma 4 causes it to malform tool calls under real load, affecting vLLM, llama.cpp, Ollama, and oobabooga. A developer open-sourced a diagnosis, repair, and experimental LoR…

00:00
2026-06-19
depot.dev
developer-tools

Now available: SOCI v2 support for Depot container builds

Depot now supports SOCI v2 for container builds, enabling lazy-pulling of images to drastically reduce startup times. The feature generates a SOCI index during the build process, allowing containers t…

18:18
2026-06-18
dev.to
developer-tools

IDE fixes, TS 5.9 beta, Claude tool use explained

The Continue plugin v1.2.20 patches memory leaks, unhandled exceptions, and JCEF message chunking crashes across JetBrains and VS Code adapters, fixing crash vectors that cause sidebar hangs and autoc…

16:29
2026-06-18
devashish.me
large-language-models

Two Qwen3 models on one DGX Spark: the residency math

Alibaba's Qwen3-80B and Qwen3-4B models were successfully co-located on a single NVIDIA DGX Spark using vLLM containers behind a LiteLLM proxy, but the 80B model's inability to emit tool calls in auto…

16:14
2026-06-18
dev.to
artificial-intelligence

7 Open-Source AI Projects Developers Need [June 2026]

Seven open-source AI projects—Ollama, Open WebUI, Browser Use, vLLM, Unsloth, CrewAI, and Continue—are reshaping production software development in June 2026. Ollama, with 174,000+ GitHub stars, now o…

10:16
2026-06-18
dev.to
large-language-models

What GLM-5.2 Changes for Long-Horizon Coding

Zhipu AI released GLM-5.2, a large language model with a 1M-token context window, flexible effort levels, and an MIT license, targeting long-horizon coding tasks. The model introduces IndexShare, an a…

09:00
2026-06-18
anyscale.com
large-language-models

High Performance Distributed Inference with Ray Serve LLM

Ray Serve LLM, in partnership with Google Kubernetes Engine, announced major performance improvements achieving up to 4.4x higher throughput on prefill-heavy workloads and 24x higher on decode-heavy w…

02:50
2026-06-18
discuss.huggingface.co
large-language-models

Local-LLM-Launcher-GUI: For those who hate CLI flags

A new open-source GUI tool, Local-LLM-Launcher-GUI, lets users run large language models locally via vLLM or llama.cpp without memorizing command-line flags. The browser-based interface provides hardw…

01:08
2026-06-18
byteiota.com
large-language-models

MiniMax M3: Open-Weight Frontier Model at 5% of Opus Cost

MiniMax released the M3 open-weight model, claiming it costs 5% of Claude Opus per task, achieves 59% on SWE-Bench Pro, and supports a 1-million-token context window at one-twentieth the compute of it…

00:00
2026-06-18
techstackups.com
large-language-models

GLM-5.2 vs Claude Opus

Z.ai released GLM-5.2, an open-weights AI model under an MIT license, positioning it between Claude Opus 4.7 and 4.8 in performance while costing less than a fifth of Opus on output tokens. The model …

← prev page 33 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics