cd/entity/SGLang· home entities SGLang
grep -l @sglang /news/*.json | wc -l → 145

SGLang

mentions 145 type Organization page 1/8 feed RSS

// recent coverage 145 mentions

21:26
2026-08-21
promptcube3.com
artificial-intelligence

Nvidia's latest demo proves the inference stack matters more

Nvidia's latest demo shows that optimizing the inference stack, not the model weights, is the key to performance, achieving 4.2× higher throughput at half the latency on identical H100 hardware with L…

19:37
2026-08-21
patronus.ai
machine-learning

Getting GLM-5.2 NVFP4 Post-Training off the ground

Z.ai and NVIDIA engineers trained GLM-5.2, a 744B-parameter mixture-of-experts model quantized to 4-bit NVFP4, with reinforcement learning to play Super Mario Bros., overcoming arithmetic, distributed…

17:36
2026-08-21
liquid.ai
artificial-intelligence

LFM2.5-DSpark: Up to 3.2x Faster Inference from H100 to MacB

Liquid AI released DSpark draft model checkpoints for three LFM2.5 models, enabling speculative decoding that speeds up inference by up to 3.18x on GPUs and 2.87x on-device without changing output qua…

16:52
2026-08-20
huggingface.co
artificial-intelligence

Up to 3.2x Faster Inference with LFM2.5-DSpark

Liquid AI released DSpark draft model checkpoints for its LFM2.5 family, claiming up to 3.18x throughput improvement on GPU and up to 2.87x on-device, with day-one support for llama.cpp and SGLang. Th…

16:08
2026-08-20
sourcefeed.dev
artificial-intelligence

You're Not Buying Compute, You're Buying Utilization

A developer's four-month home-lab test found that self-hosting open models costs about 7× more than using hosted APIs, cutting their bill from roughly $3,050 to $420 per month by switching to API acce…

00:00
2026-08-20
mindstudio.ai
large-language-models

How to Run Qwen 3 27B Locally with DeepSeek Harness

Alibaba's Qwen 3 27B, a dense open-weight multimodal model, can be run locally using DeepSeek Harness, achieving agentic performance close to Claude 4.5 on benchmarks. On a two-node NVIDIA DGX Spark c…

14:30
2026-08-19
hiraditya.github.io
artificial-intelligence

Two Schedulers, One SLO

A vLLM RFC from the llm-d team warns that disaggregated inference deployments, where prefill and decode run on separate schedulers, can trigger recomputation-based preemption inside the decode instanc…

00:20
2026-08-19
inco.ai
artificial-intelligence

DFlash 2: Keep Drafting Parallel

Inco AI released DFlash 2, a parallel speculative decoding technique that delivers over 20% more output from every verification pass with around 1% added cycle latency, achieving 2.7–3.4× throughput o…

23:21
2026-08-18
baseten.co
ai-infrastructure

Inference Engineering by Philip Kiely – Digital Download

Philip Kiely's new book, 'Inference Engineering,' is now available as a digital download, offering a comprehensive guide to the technologies and techniques powering AI inference across runtime, infras…

19:22
2026-08-18
runtimewire.com
artificial-intelligence

RadixArk released Miles v0.1 to align AI training and inference

RadixArk released Miles v0.1, an open-source framework for aligning AI training and inference, on GitHub. The company, founded by Ying Sheng and Banghua Zhu, aims to commercialize open-source infrastr…

18:02
2026-08-18
lmsys.org
machine-learning

Miles v0.1: Production-level Post-training

Radix Ark released Miles v0.1, a full-stack production-ready system for frontier post-training, designed to make large-scale reinforcement learning accessible to researchers and developers. The system…

00:00
2026-08-17
rocm.blogs.amd.com
ai-tools

Bring Claude Code On‑Prem with AMD Instinct GPUs

Anthropic's Claude Code can now run on-premises with AMD Instinct GPUs, serving GLM 5.2 at full quality via SGLang and LiteLLM, eliminating cloud API dependency and per-token costs. The setup uses an …

16:21
2026-08-15
github.com
ai-tools

Qwen 3.8 27B at 2x speed on a 5090

Adore LLC released Balto Speedrunner, a Windows app that turns Qwen 3.8 27B into a local coding agent for a single RTX 5090, achieving up to 300 tok/s on chat prompts and 150 tok/s during coding runs.…

07:12
2026-08-15
byteiota.com
artificial-intelligence

Qwen3.8-27B Is Out: The Local AI Model Developers Need

Alibaba released Qwen3.8-27B, a 27.78-billion-parameter open-weights multimodal model under Apache 2.0, with a 262,144-token native context window and configurable reasoning. Vendor-reported benchmark…

page 1 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics