cd/entity/vLLM· home entities vLLM
grep -l @vllm /news/*.json | wc -l → 443

vLLM

mentions 443 type Organization page 9/23 feed RSS

// recent coverage 443 mentions

15:13
2026-07-27
huggingface.co
artificial-intelligence

Kimi K3 License

Moonshot AI has released Kimi-K3, an image-text-to-text model under a permissive license that allows use, modification, and commercial distribution, with support for Transformers, vLLM, SGLang, and Do…

12:15
2026-07-27
byteiota.com
large-language-models

Kimi K3 Open Weights Live: Self-Host or Use the API?

Moonshot AI released the open weights of its 2.8-trillion-parameter Kimi K3 sparse Mixture-of-Experts model on Hugging Face under an Apache 2.0 license, but the 594 GB MXFP4-quantized model requires a…

12:08
2026-07-27
sourcefeed.dev
artificial-intelligence

The Real Cost of 'Just Use vLLM,' According to Netflix

Netflix's AI Platform team published a detailed account of its LLM serving platform, revealing that version pinning between NVIDIA Triton Inference Server and vLLM, a Python GIL bottleneck, and KV-cac…

00:02
2026-07-27
promptcube3.com
artificial-intelligence

Prefill-Decode Disaggregation

Prefill-decode disaggregation separates the compute-bound prefill and memory-bound decode phases of LLM inference onto different hardware to solve the 'noisy neighbor' problem, according to a technica…

17:41
2026-07-26
sourcefeed.dev
artificial-intelligence

Autoscale GPU Inference on EKS with Karpenter and Spot Instances

Karpenter v1.14.0 can autoscale GPU inference on Amazon EKS by provisioning spot GPU nodes on demand, bin-packing a vLLM v0.25.1 model server, and deleting nodes when traffic drops, eliminating static…

13:46
2026-07-26
promptcube3.com
large-language-models

Local LLM Deployment: A Practical Guide

Deploying large language models locally requires matching hardware to model size, with quantization enabling massive models to run on consumer hardware. Ollama, LM Studio, and vLLM are recommended too…

05:18
2026-07-26
kraghavan.ca
large-language-models

Introduction to LLM Inference

A senior engineer with 11 years of distributed systems experience explains the full LLM inference pipeline, from request arrival to text output, detailing the GGUF file structure and the distinction b…

03:02
2026-07-26
promptcube3.com
artificial-intelligence

OpenAI Models Leaking to Hugging Face: Analysis

OpenAI models are appearing on Hugging Face, offering researchers a rare opportunity to audit proprietary architectures through leaked weights, config files, and inference testing. The leaks expose hi…

19:46
2026-07-25
promptcube3.com
large-language-models

DeepSeek-R1 Local Deployment: My Hardware Struggles

A user reports that deploying the full 671B parameter DeepSeek-R1 model locally requires over 100GB of VRAM and is impractical on consumer hardware, with CUDA out-of-memory errors occurring even at sm…

17:03
2026-07-25
promptcube3.com
artificial-intelligence

Open-Weight AI: Model Wars vs Ecosystem Wars

Open-weight AI models offer freedom but require significant effort to deploy, according to a technical guide that argues the real value lies in deployment pipelines and developer ecosystems rather tha…

14:47
2026-07-25
promptcube3.com
large-language-models

KV Caching: Why Your LLM Inference Costs are Sky-High

KV caching, which stores Key and Value tensors in GPU memory to avoid recalculating attention for every token, is the primary driver of high LLM inference costs because the cache grows linearly with s…

19:03
2026-07-24
promptcube3.com
artificial-intelligence

Open Source AI: Why Closed-Source Lobbying is Failing

Open-source AI is winning over closed-source lobbying due to deployment flexibility, cost, and hardware ecosystem scale, according to a developer analysis. The global infrastructure for running open w…

16:09
2026-07-24
developers.googleblog.com
artificial-intelligence

Run Ray on TPU, Part 2: Ray AI libraries

Ray AI libraries (Serve, Data, Train) now support Google TPU slices through a topology field that reserves a whole ICI-connected slice, preventing multi-host deployment hangs. Ray Serve serves LLMs vi…

← prev page 9 / 23 next →
// co-occurs with top 8 entities
// topics top 6 topics