cd/entity/SGLang· home› entities› SGLang
grep -l @sglang /news/*.json | wc -l → 236

SGLang

mentions 236 type Organization page 2/12 feed RSS

// recent coverage 236 mentions

00:00
2026-09-22
rocm.blogs.amd.com
large-language-models

Benchmarking Kimi-K3 Across vLLM, SGLang, and ATOM on MI350X

Moonshot AI released the weights for Kimi-K3, a 2.8-trillion-parameter, 1M-context, natively-MXFP4 Mixture-of-Experts model, with AMD Instinct support available on day 0 across three serving framework…

04:05
2026-09-21
tokenstead.ai
generative-ai

Qwen-Image-2.1

Alibaba's Qwen released Qwen-Image-2.1, a 7B-parameter image generation and editing model that combines a 32-layer single-stream DiT with a Qwen3-VL 8B text encoder and a 64-channel RGBA VAE, but ship…

00:00
2026-09-21
simm.is
ai-infrastructure

A KV cache you can fork

Replikativ released pretrained-rstr, an MIT-licensed inference component that turns a language model's KV cache into a forkable, content-addressed value rather than a process-local optimization. The s…

04:07
2026-09-20
arxiv.org
ai-safety

Inference-Engine Fingerprinting Attacks Are Practical

A September 17, 2026 arXiv paper shows that a misaligned AI model can fingerprint which inference engine executes it — including vLLM and SGLang — and then use engine-specific exploits to seize contro…

00:25
2026-09-17
runtimewire.com
ai-research

Periodic Labs details the 1,300-GPU stack behind Neon

Periodic Labs disclosed in a September 15th engineering post that its Neon 1-trillion-parameter model's final training run peaked at 1,300 Nvidia H200 GPUs, achieving over 95% cluster utilization, 4.1…

00:00
2026-09-17
signoz.io
ai-infrastructure

SGLang Monitoring and Observability with OpenTelemetry

SigNoz published a guide for monitoring self-hosted SGLang inference servers by sending Prometheus metrics and OpenTelemetry request traces to its observability platform. The setup requires launching …

23:16
2026-09-16
frontierroles.com
ai-research

Research, Tinker, RL Systems — Thinking Machines Lab

Thinking Machines Lab is hiring a Research, Tinker, RL Systems engineer in San Francisco at an annual salary range of $350,000 to $475,000, a figure the job board states sits 73% above the $238,000 me…

04:51
2026-09-16
lmsys.org
ai-infrastructure

SGLang and Miles Add Day-0 Support for DeepSeek-v4.1

SGLang and Miles added day-0 support for DeepSeek-V4.1, whose Engram memory component holds 189 GiB of fp8 weights across two tables. In paired tests on 4x GB300 (TP4/EP4), host offload of the Engram …

20:55
2026-09-15
superml.dev
ai-infrastructure

Prefill, Not Decode, Is Your Agent's Real Bottleneck

Prefill, not decode, has become the dominant bottleneck in RAG and multi-agent workloads, according to an analysis citing NVIDIA's published figures of roughly 30x higher served-request counts for lar…

16:00
2026-09-15
gladlabs.io
ai-infrastructure

Llama.cpp vs vLLM vs SGLang

Glad Labs decided not to switch its self-hosted inference stack from Ollama to vLLM after reviewing its own call logs, which showed only one to three concurrent calls at most against roughly 50 calls …

12:09
2026-09-13
sourcefeed.dev
ai-research

Transformers v5 turned a library into a standard

Hugging Face released Transformers v5.17.0 on September 9, adding seven model architectures in a single minor release, including Tencent's 780-billion-parameter mixture-of-experts model and Moonshot A…

00:00
2026-09-13
mindstudio.ai
ai-products

How to Self-Host Nex-N2.5 with SGLang and Docker

Nex-AGI released its Nex-N2.5 family of open-weight agentic models in three sizes — mini, Pro, and Max — with a prebuilt Docker image running a customized SGLang fork (nexagi/sglang:v0.5.18-nex-patch)…

← prev page 2 / 12 next →
// co-occurs with top 8 entities
// topics top 6 topics