cd/entity/SGLang· home entities SGLang
grep -l @sglang /news/*.json | wc -l → 145

SGLang

mentions 145 type Organization page 3/8 feed RSS

// recent coverage 145 mentions

19:02
2026-08-05
tokenstead.ai
large-language-models

Pokee-Isaac 28B

Pokee AI released Pokee-Isaac 28B, a 28B-parameter proprietary non-decoder-only model claiming a 10M-token context that fits on a single RTX 4090 (24GB) in quantized form. Vendor-reported benchmarks i…

02:49
2026-08-05
dev.to
large-language-models

Measuring LLM Prefix Caching: The Cache Hit Rate Metric

An engineer's benchmarking guide introduces a cache hit rate metric for measuring prefix caching effectiveness in LLM serving, implemented in the open-source tool llmperf-rs. The metric calculates the…

23:08
2026-08-03
sourcefeed.dev
artificial-intelligence

Cloudflare's Quantization Math Is Right. The Disclosure Isn't

Cloudflare published an engineering post detailing how it quantizes models on Workers AI, including Moonshot's Kimi K2.6 with FP8 KV cache and Z.ai's GLM 5.2 with INT4 weights, achieving up to 41% hig…

21:09
2026-08-03
sourcefeed.dev
artificial-intelligence

Quantize the Decode, Not the Prefill

Cloudflare's production benchmarks for Moonshot's Kimi K2.6 and Z.ai's GLM 5.2 show that quantizing the decode phase, not the prefill, yields the biggest cost and throughput gains for trillion-paramet…

13:00
2026-08-03
blog.cloudflare.com
machine-learning

Smaller, faster, safer: running Kimi and GLM at scale

Cloudflare's Workers AI has implemented three optimizations—quantizing the KV cache, compressing model weights, and protecting the shared cache—to run Moonshot's Kimi K2.6 and Z.ai's GLM 5.2 more effi…

13:59
2026-07-30
openalternative.co
artificial-intelligence

LocalAI

LocalAI, a self-hosted runtime that runs AI workloads on user-controlled hardware, offers an OpenAI-compatible API supporting text generation, vision, speech, image and video generation, embeddings, a…

09:21
2026-07-29
blog.us.fixstars.com
large-language-models

Can a 2.8T Model Run on a Single Node of Nvidia B300 X8?

Moonshot AI released the 2.8-trillion-parameter Kimi-K3 open-weight model on July 27, 2026, and Fixstars successfully ran inference on a single-node NVIDIA B300 x8 system using SGLang's DCP support, a…

09:47
2026-07-28
comfyfile.com
large-language-models

How to Download and Run Kimi K3 Open Weights

Moonshot AI released the full Kimi K3 open weights on July 27, 2026, a 2.8-trillion-parameter mixture-of-experts model with a one-million-token context window and a 1.4 TB download. The model uses nat…

17:23
2026-07-27
cryptobriefing.com
artificial-intelligence

Moonshot completes Kimi K3 rollout with full model weight release

Moonshot AI released the full Kimi K3 model weights and technical report on July 27, 2026, making its 2.8 trillion parameter mixture-of-experts model available to developers and researchers. The model…

15:16
2026-07-27
baseten.co
artificial-intelligence

We built a day-0 API for Kimi K3

Baseten has launched day-0 API support for Kimi K3, a new open frontier model from Moonshot AI with 2.8 trillion parameters, making it the largest open model to date. The API runs on NVIDIA GB300 NVL7…

15:13
2026-07-27
huggingface.co
artificial-intelligence

Kimi K3 License

Moonshot AI has released Kimi-K3, an image-text-to-text model under a permissive license that allows use, modification, and commercial distribution, with support for Transformers, vLLM, SGLang, and Do…

← prev page 3 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics