cd/entity/SGLang· home entities SGLang
grep -l @sglang /news/*.json | wc -l → 146

SGLang

mentions 146 type Organization page 6/8 feed RSS

// recent coverage 146 mentions

15:27
2026-06-27
cefboud.com
large-language-models

Distributed LLM Inference with LLM-d

A new open-source tool called llm-d acts as an LLM-aware load balancer for distributed inference, intelligently routing requests across vLLM instances based on KV cache locality and GPU utilization. B…

20:51
2026-06-24
blog.crossplane.io
ai-infrastructure

I built a fleet-scale inference control plane using Crossplane

A developer built Modelplane, an open-source inference control plane using Crossplane, to manage GPU fleets across clouds, neoclouds, and on-premise environments. The platform allows platform teams to…

18:00
2026-06-23
research.ibm.com
artificial-intelligence

Running AI on mixed hardware for speed and affordability

IBM Research, Red Hat, and NxtGen Cloud Technologies demonstrated that using llm-d to serve AI models on mixed GPU hardware can boost inference speeds by 3 to 5 times and double throughput, enabling e…

11:35
2026-06-23
github.com
artificial-intelligence

Unlimited OCR: One-shot long-horizon parsing

Baidu released Unlimited-OCR, a one-shot long-horizon parsing model that extends DeepSeek-OCR, on June 22, 2026. The open-source model supports single-image and multi-page PDF parsing with configurabl…

00:00
2026-06-23
modelplane.ai
ai-infrastructure

Introducing Modelplane: the control plane for AI inference

Modelplane, an open-source control plane for AI inference built on Crossplane, is being released to manage GPU clusters as a single inference fleet, handling provisioning, model placement, autoscaling…

09:22
2026-06-22
blog.doubleword.ai
large-language-models

FlashOffload: 7x Cheaper Prefills with Offloading

Researchers improved SGLang's offloading engine to achieve 7x cheaper prefill costs for DeepSeek V4 Flash on Grace Hopper systems, leveraging high CPU-GPU bandwidth to hide weight transfers behind com…

00:17
2026-06-20
modal.com
large-language-models

Speculation Is All You Need

Modal Labs released state-of-the-art DFlash speculators for Qwen 3.5 and Qwen 3.6 models on Hugging Face, achieving 5-20% additional speedups and enabling Qwen 3.5 122B-A10B to run at over 1000 tok/s …

20:25
2026-06-19
lmsys.org
large-language-models

The next generation of speculative decoding: DFlash and Spec V2

Modal and Z Lab released DFlash, a speculative decoding model for Qwen 3.5 397B-A17B, achieving over 4.3x throughput versus baseline and 1.5x versus MTP on HumanEval at concurrency 1. The model uses a…

10:16
2026-06-18
dev.to
large-language-models

What GLM-5.2 Changes for Long-Horizon Coding

Zhipu AI released GLM-5.2, a large language model with a 1M-token context window, flexible effort levels, and an MIT license, targeting long-horizon coding tasks. The model introduces IndexShare, an a…

← prev page 6 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics