cd/entity/vLLM· home› entities› vLLM
grep -l @vllm /news/*.json | wc -l → 748

vLLM

mentions 748 type Organization page 24/38 feed RSS

// recent coverage 748 mentions

09:47
2026-07-28
comfyfile.com
large-language-models

How to Download and Run Kimi K3 Open Weights

Moonshot AI released the full Kimi K3 open weights on July 27, 2026, a 2.8-trillion-parameter mixture-of-experts model with a one-million-token context window and a 1.4 TB download. The model uses nat…

19:13
2026-07-27
promptcube3.com
large-language-models

Kimi K3: Why Open-Weight Models are Shaking Up the Market

Open-weight AI models like Kimi K3 are disrupting the market by enabling local deployment and customization without proprietary fees, according to the article. The shift lowers barriers for developers…

17:23
2026-07-27
cryptobriefing.com
artificial-intelligence

Moonshot completes Kimi K3 rollout with full model weight release

Moonshot AI released the full Kimi K3 model weights and technical report on July 27, 2026, making its 2.8 trillion parameter mixture-of-experts model available to developers and researchers. The model…

15:44
2026-07-27
vllm.ai
artificial-intelligence

Kimi K3 on vLLM: Up to 370 Tokens/sec

VLLM announces efficient day-0 support for Moonshot AI's Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model, achieving up to 370 tokens per second with speculative decoding on 16 NVIDIA GB300 …

15:16
2026-07-27
baseten.co
artificial-intelligence

We built a day-0 API for Kimi K3

Baseten has launched day-0 API support for Kimi K3, a new open frontier model from Moonshot AI with 2.8 trillion parameters, making it the largest open model to date. The API runs on NVIDIA GB300 NVL7…

15:13
2026-07-27
huggingface.co
artificial-intelligence

Kimi K3 License

Moonshot AI has released Kimi-K3, an image-text-to-text model under a permissive license that allows use, modification, and commercial distribution, with support for Transformers, vLLM, SGLang, and Do…

12:15
2026-07-27
byteiota.com
large-language-models

Kimi K3 Open Weights Live: Self-Host or Use the API?

Moonshot AI released the open weights of its 2.8-trillion-parameter Kimi K3 sparse Mixture-of-Experts model on Hugging Face under an Apache 2.0 license, but the 594 GB MXFP4-quantized model requires a…

12:08
2026-07-27
sourcefeed.dev
artificial-intelligence

The Real Cost of 'Just Use vLLM,' According to Netflix

Netflix's AI Platform team published a detailed account of its LLM serving platform, revealing that version pinning between NVIDIA Triton Inference Server and vLLM, a Python GIL bottleneck, and KV-cac…

00:02
2026-07-27
promptcube3.com
artificial-intelligence

Prefill-Decode Disaggregation

Prefill-decode disaggregation separates the compute-bound prefill and memory-bound decode phases of LLM inference onto different hardware to solve the 'noisy neighbor' problem, according to a technica…

17:41
2026-07-26
sourcefeed.dev
artificial-intelligence

Autoscale GPU Inference on EKS with Karpenter and Spot Instances

Karpenter v1.14.0 can autoscale GPU inference on Amazon EKS by provisioning spot GPU nodes on demand, bin-packing a vLLM v0.25.1 model server, and deleting nodes when traffic drops, eliminating static…

13:46
2026-07-26
promptcube3.com
large-language-models

Local LLM Deployment: A Practical Guide

Deploying large language models locally requires matching hardware to model size, with quantization enabling massive models to run on consumer hardware. Ollama, LM Studio, and vLLM are recommended too…

05:18
2026-07-26
kraghavan.ca
large-language-models

Introduction to LLM Inference

A senior engineer with 11 years of distributed systems experience explains the full LLM inference pipeline, from request arrival to text output, detailing the GGUF file structure and the distinction b…

← prev page 24 / 38 next →
// co-occurs with top 8 entities
// topics top 6 topics