cd/entity/SGLang· home entities SGLang
grep -l @sglang /news/*.json | wc -l → 145

SGLang

mentions 145 type Organization page 5/8 feed RSS

// recent coverage 145 mentions

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

GEAK Agent-Driven Optimization of the DeepSeekV4 MLA Kernel

AMD's open-source GEAK agent-driven framework automated the optimization of the DeepSeekV4 MLA kernel, achieving a 2.10x improvement in end-to-end throughput and a 3.71x reduction in time-to-first-tok…

00:00
2026-07-13
rocm.blogs.amd.com
artificial-intelligence

QuickReduce INT3 Quantization and Benchmarking on MI355

AMD's QuickReduce library now supports INT3 quantization for all-reduce communication in multi-GPU LLM inference, achieving a 22% reduction in on-wire data volume compared to INT4 on AMD Instinct MI35…

15:21
2026-07-10
byteiota.com
artificial-intelligence

NVIDIA Nemotron-Labs-Diffusion Kills the Draft Model

NVIDIA released Nemotron-Labs-Diffusion, a single model that eliminates the need for a separate draft model in speculative decoding, achieving 6.82 accepted tokens per forward pass in self-speculation…

07:10
2026-07-08
byteiota.com
artificial-intelligence

Tencent Hy3: 295B MoE Hits SWE-Bench 78 — Free API Ends July 21

Tencent released Hy3, a 295B Mixture-of-Experts model with Apache 2.0 weights, achieving a 78.0 SWE-bench Verified score and leading open-weight models on tool and search benchmarks. A free API on Ope…

00:00
2026-07-08
rocm.blogs.amd.com
artificial-intelligence

SGLang-ATOM: Bring ROCm-Native Acceleration to SGLang Serving

AMD introduced SGLang-ATOM, a bridge connecting the SGLang serving framework with ATOM's ROCm-native execution path to accelerate large language model inference on AMD Instinct GPUs. The integration u…

15:20
2026-07-07
huggingface.co
artificial-intelligence

Hugging Face Models on Foundry Managed Compute

Microsoft Foundry now offers a curated catalog of Hugging Face open-weight models deployable on Foundry Managed Compute, with pre-staged weights in Azure and built-in enterprise security, governance, …

03:01
2026-07-07
lmsys.org
ai-agents

Agent-Assisted SGLang Development: An Initial Exploration

SGLang development is being augmented with agent-assisted workflows that encode procedural engineering knowledge into executable skills, covering LLM serving, GPU kernels, diffusion pipelines, and pro…

00:42
2026-07-07
letsdatascience.com
large-language-models

Tencent open-sources Hy3 295B MoE model

Tencent released Hy3, an Apache-2.0 open-weight Mixture-of-Experts model with 295B total parameters, 21B active parameters per token, and a 256K context window. The release includes official artifacts…

21:25
2026-07-02
developer.nvidia.com
ai-safety

Hardware-Rooted AI Security That Won’t Slow You Down

NVIDIA announced that its Confidential Computing technology for Blackwell GPUs achieves up to 98% of the inference performance of non-secure solutions, enabling hardware-rooted AI security without sig…

09:07
2026-07-01
glukhov.org
large-language-models

Speculative Decoding: 20-50% Faster LLM Inference

Speculative decoding accelerates large language model inference by 20-50% without quality loss, using a draft-verify mechanism that generates multiple tokens per forward pass. The technique amortizes …

03:09
2026-07-01
byteiota.com
large-language-models

MiniMax M3: Open-Weight Model That Beats GPT-5.5 on Coding

MiniMax released M3, a 428-billion-parameter open-weight model, on June 7, achieving 59.0% on SWE-Bench Pro—slightly outperforming GPT-5.5's 58.6%—at $0.30 per million input tokens, making it 16 times…

← prev page 5 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics