cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 3/23 feed RSS

// recent coverage 443 mentions

20:43
2026-08-14
dev.to
developer-tools

Serving Gemma4 with Rust on vLLM 🦀

A developer detailed how to build and run vLLM's Rust frontend on an AWS EC2 G5g instance with Graviton2 and an NVIDIA T4G GPU. The tutorial highlights that vLLM now requires a Rust toolchain for sour…

16:45
2026-08-14
promptcube3.com
artificial-intelligence

vLLM beats Ollama by 20x once you hit high concurrency

VLLM outperforms Ollama by nearly 20x in throughput at high concurrency, according to benchmark tests running Llama 3.1 8B on an NVIDIA A100 40GB, with vLLM peaking at 793 tokens per second versus Oll…

12:16
2026-08-14
dev.to
large-language-models

vLLM vs Ollama: Production Serving 2026

A developer's comparison of vLLM and Ollama for LLM serving in 2026 shows that while Ollama is simpler for single-user local use, vLLM outperforms it dramatically under concurrency, with throughput up…

06:16
2026-08-14
latent.space
artificial-intelligence

[AINews] Cursor's $60B acquisition by SpaceXai closes

Z.ai launched GLM-5.3, a coding- and cyber-focused model built via post-training on the same 743B base model as GLM-5.2, achieving scores of 28.3 on Terminal Bench 3.0, 66.9 on DeepSWE, 28.5 on Agents…

00:00
2026-08-14
mindstudio.ai
large-language-models

How to Run DeepSeek V4 Pro Locally with vLLM or SGLang

DeepSeek released DeepSeek-V4-Pro-0813, the production successor to its V4 Pro preview, under an MIT license with open weights, adding a DSpark speculative decoding module and stronger agentic benchma…

20:01
2026-08-13
pub.towardsai.net
large-language-models

Start Here: The Words Everyone Uses About LLM Inference

In a new series on LLM inference, the author explains the core concepts behind running language models in production, starting with the fundamental division between prefill and decode. The series cove…

18:49
2026-08-13
acefleet.dev
ai-infrastructure

Scale Your AI Revenue – Not Your Cloud Bill

A new guide outlines strategies for scaling AI revenue while controlling cloud costs, covering hardware accelerators from NVIDIA, Groq, and Cerebras, inference engines like vLLM and TensorRT-LLM, and …

18:45
2026-08-13
dev.to
large-language-models

Running Gemma 4 on EC2 G5g: Graviton2 AMD with NVIDIA GPU

An engineer successfully ran Google's Gemma 4 E2B model on AWS EC2 G5g, a Graviton2 (aarch64) instance with an NVIDIA T4G GPU, achieving 43.1 tokens per second after patching vLLM. The deployment requ…

11:01
2026-08-13
dev.to
large-language-models

ChatGPT Desktop for Linux: A new way to interact!

A developer detailed the architectural requirements for building a Linux-native desktop client for LLM-based coding assistants, emphasizing the need for multi-process designs, Rust backends, and deep …

02:33
2026-08-13
github.com
ai-products

Doable: Self-hosted AI app builder for teams

Doable, a self-hosted AI app builder for teams, has been released under an MIT license, allowing users to generate, deploy, and host AI-powered applications on their own infrastructure with multi-tena…

← prev page 3 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics