How to Download and Run Kimi K3 Open Weights
Moonshot AI released the full Kimi K3 open weights on July 27, 2026, a 2.8-trillion-parameter mixture-of-experts model with a one-million-token context window and a 1.4 TB download. The model uses nat…
Moonshot AI released the full Kimi K3 open weights on July 27, 2026, a 2.8-trillion-parameter mixture-of-experts model with a one-million-token context window and a 1.4 TB download. The model uses nat…
Moonshot AI released Kimi K3, an open-source large language model with 2.8 trillion parameters and a 1-million-token context window, in July 2026. It is the first open-source model to surpass Claude a…
LOCKS, a new method from arXiv, enables efficient long-context decoding by giving each page of the key-value cache its own compact spectral summary, reconstructing within-page logits, and attending on…
Open-weight AI models like Kimi K3 are disrupting the market by enabling local deployment and customization without proprietary fees, according to the article. The shift lowers barriers for developers…
Moonshot AI released the full Kimi K3 model weights and technical report on July 27, 2026, making its 2.8 trillion parameter mixture-of-experts model available to developers and researchers. The model…
VLLM announces efficient day-0 support for Moonshot AI's Kimi K3, a 2.8-trillion-parameter Mixture-of-Experts model, achieving up to 370 tokens per second with speculative decoding on 16 NVIDIA GB300 …
Baseten has launched day-0 API support for Kimi K3, a new open frontier model from Moonshot AI with 2.8 trillion parameters, making it the largest open model to date. The API runs on NVIDIA GB300 NVL7…
Moonshot AI has released Kimi-K3, an image-text-to-text model under a permissive license that allows use, modification, and commercial distribution, with support for Transformers, vLLM, SGLang, and Do…
Moonshot AI released the open weights of its 2.8-trillion-parameter Kimi K3 sparse Mixture-of-Experts model on Hugging Face under an Apache 2.0 license, but the 594 GB MXFP4-quantized model requires a…
Netflix's AI Platform team published a detailed account of its LLM serving platform, revealing that version pinning between NVIDIA Triton Inference Server and vLLM, a Python GIL bottleneck, and KV-cac…
Netflix detailed its in-house LLM serving platform built on Triton and vLLM, revealing how it handles real-time and batch inference across CPUs and GPUs while managing version compatibility and constr…
Prefill-decode disaggregation separates the compute-bound prefill and memory-bound decode phases of LLM inference onto different hardware to solve the 'noisy neighbor' problem, according to a technica…
Prefill-decode disaggregation, now supported by vLLM, SGLang, LLM-d, NVIDIA Dynamo, and TensorRT-LLM, separates LLM inference into compute-bound prefill and memory-bound decode phases on different nod…
LightSeek Org released Shepherd Model Gateway (SMG), an engine-agnostic, high-performance model-routing gateway for large-scale LLM deployments that centralizes worker lifecycle management and balance…
Karpenter v1.14.0 can autoscale GPU inference on Amazon EKS by provisioning spot GPU nodes on demand, bin-packing a vLLM v0.25.1 model server, and deleting nodes when traffic drops, eliminating static…
Deploying large language models locally requires matching hardware to model size, with quantization enabling massive models to run on consumer hardware. Ollama, LM Studio, and vLLM are recommended too…
OpenLake, an open-source storage engine for offloading LLM KV caches from GPU memory to RAM and NVMe, cuts GPU time by 48.2% for long-context inference, reducing a 1,169-second workload to 606 seconds…
ARIA, a voice-native 3D spatial AI security operations cockpit with governed autonomy, is now available under BSL 1.1 for evaluation and research. Developed by a solo developer, the platform runs enti…
A developer built a local-first voice-enabled AI assistant by combining Nous Research's open-source Hermes Agent framework with Kokoro TTS, achieving natural speech responses without cloud API costs o…
A senior engineer with 11 years of distributed systems experience explains the full LLM inference pipeline, from request arrival to text output, detailing the GGUF file structure and the distinction b…