cd/entity/SGLang· home entities SGLang
grep -l @sglang /news/*.json | wc -l → 145

SGLang

mentions 145 type Organization page 2/8 feed RSS

// recent coverage 145 mentions

17:56
2026-08-14
twitter.com
large-language-models

Qwen3.8 27B Runs 200 tok/s on a Single RTX 5090

Alibaba's Qwen3.8-27B open-source model achieves 206.1 tokens per second decode on a single Nvidia RTX 5090 using SGLang's NVFP4 and DSpark, with Day-0 support in SGLang. The 27B-parameter multimodal …

00:00
2026-08-14
mindstudio.ai
large-language-models

How to Run DeepSeek V4 Pro Locally with vLLM or SGLang

DeepSeek released DeepSeek-V4-Pro-0813, the production successor to its V4 Pro preview, under an MIT license with open weights, adding a DSpark speculative decoding module and stronger agentic benchma…

11:02
2026-08-11
blog.doubleword.ai
ai-infrastructure

On-the-fly snapshot compression for elastic inference at scale

CRIU now supports LZ4 compression directly in its memory checkpoint and restore paths, enabling on-the-fly snapshot compression for elastic inference at scale. The feature compresses memory pages duri…

00:00
2026-08-11
modelplane.ai
artificial-intelligence

Why Day 0 for Nemotron 3.5 Lightning wasn't a scramble

NVIDIA released Nemotron-3.5-Lightning, a 30B mixture-of-experts model with 3B active parameters, on the same day Modelplane, an open-source fleet-level control plane for inference, achieved zero-day …

19:19
2026-08-10
runtimewire.com
ai-agents

Cloudflare absolutely crushed its five-day Agents Week sprint

Cloudflare capped its five-day Agents Week sprint on August 7, 2026, with more than two dozen announcements aimed at positioning itself as the operating layer for software agents, integrating its Work…

11:49
2026-08-10
dev.to
artificial-intelligence

Only Two AI Updates Cleared My 36-Hour Cutoff

A developer's review of recent AI releases found only two updates met a strict 36-hour cutoff: Meta's Muse Glimmer, a 30B multimodal model under Apache 2.0, and Hugging Face's Transformers 5.15.0. The…

10:31
2026-08-10
twitter.com
artificial-intelligence

Muse Spark 1.2 (Meta) is reportedly becoming open-weight

Meta announced it will release an open-weight version of Muse Spark 1.2 and released Muse Glimmer, a 30B agentic model with open weights under Apache 2.0, which can run on 24GB of VRAM. The model is q…

08:18
2026-08-10
ssenthilnathan3.github.io
machine-learning

MoE routing is just branch prediction

A software engineer's analysis argues that MoE routing in transformer inference is fundamentally the same problem as CPU branch prediction, and that KV cache management techniques such as prefix cachi…

← prev page 2 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics