cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 15/38 feed RSS

// recent coverage 748 mentions

15:31
2026-08-22
blog.bytebytego.com
large-language-models

EP223: Ollama vs vLLM vs SGLang

ByteByteGo's EP223 newsletter compares Ollama, vLLM, and SGLang for serving open-weight models, noting Ollama suits local development, vLLM handles high-traffic serving with continuous batching and Pa…

21:26
2026-08-21
promptcube3.com
artificial-intelligence

Nvidia's latest demo proves the inference stack matters more

Nvidia's latest demo shows that optimizing the inference stack, not the model weights, is the key to performance, achieving 4.2× higher throughput at half the latency on identical H100 hardware with L…

12:21
2026-08-21
matthusby.github.io
artificial-intelligence

Real world(ish) DeepSeek V4 Flash performance on a single MI300X

A single AMD MI300X GPU can serve 32 concurrent coding agents running DeepSeek V4 Flash, delivering 582 generation tokens per second after tuning, according to a benchmark by developer Ryan Zhou. The …

19:57
2026-08-20
forum.level1techs.com
artificial-intelligence

DeepSeek V4 Flash on 8× AMD gfx1201: packaged TP=8 deployment

DeepSeek V4 Flash, a 284B-parameter mixture-of-experts model with 256 routed experts and FP4 expert weights, was successfully deployed on eight AMD Radeon AI PRO R9600D GPUs (32 GB each, 256 GB total)…

19:53
2026-08-20
frontierroles.com
artificial-intelligence

Research Engineer - Agent Memory — Mem0

Mem0, a startup building long-term memory for AI agents, is hiring a Research Engineer for Agent Memory in San Francisco with a salary of $175k–250k/yr. The role involves fine-tuning models for memory…

18:00
2026-08-20
dev.to
large-language-models

Self-Hosting Kimi K3: Hardware, Cost and Sovereignty

Moonshot AI released the weights for its 2.8-trillion-parameter Kimi K3 model on 27 July 2026, enabling self-hosting but requiring at least 1.7 TB of VRAM, which rules out standard eight-way H100 node…

16:08
2026-08-20
sourcefeed.dev
artificial-intelligence

You're Not Buying Compute, You're Buying Utilization

A developer's four-month home-lab test found that self-hosting open models costs about 7× more than using hosted APIs, cutting their bill from roughly $3,050 to $420 per month by switching to API acce…

00:00
2026-08-20
runagentrun.co.uk
large-language-models

Dual 3090s: the bottleneck isn't the GPU

Two benchmarks of Qwen3.8-27B on a single RTX 3090 show a 3.2x performance gap: 41.49 tok/s with llama.cpp (build b10088) versus 132 tok/s with vLLM using a DFlash2 block drafter, according to Insider…

00:00
2026-08-20
mindstudio.ai
artificial-intelligence

Ornith 1.5 35B-A3B: Local Deployment, VRAM, and Real-World Tests

Deep Recurse released Ornith 1.5 35B-A3B, a mixture-of-experts language model with 35 billion total parameters and 3 billion active per token, built on a Qwen3.5-MoE architecture for coding, reasoning…

00:00
2026-08-20
wpnews
artificial-intelligence

Measuring Agentgateway's Overhead as an Inference Gateway

A Google Summer of Code 2026 project by Abhay Chaurasiya, mentored by Nina Polshakova and Daneyon Hansen of CNCF, measured the overhead of agentgateway as an inference gateway and found it dramaticall…

00:00
2026-08-20
mindstudio.ai
artificial-intelligence

Ornith 1.5 9B: Local Test Results Expose a Benchmark Gap

Ornith 1.5 9B, a 9-billion-parameter dense language model from ornith-ai, failed real-world agentic and coding tasks in a local test on a single Nvidia A100 GPU, despite benchmark claims that it rival…

00:00
2026-08-20
mindstudio.ai
large-language-models

How to Run Qwen 3 27B Locally with DeepSeek Harness

Alibaba's Qwen 3 27B, a dense open-weight multimodal model, can be run locally using DeepSeek Harness, achieving agentic performance close to Claude 4.5 on benchmarks. On a two-node NVIDIA DGX Spark c…

← prev page 15 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics