cd/entity/SGLang· home entities SGLang
grep -l @sglang /news/*.json | wc -l → 146

SGLang

mentions 146 type Organization page 7/8 feed RSS

// recent coverage 146 mentions

02:50
2026-06-18
discuss.huggingface.co
large-language-models

Local-LLM-Launcher-GUI: For those who hate CLI flags

A new open-source GUI tool, Local-LLM-Launcher-GUI, lets users run large language models locally via vLLM or llama.cpp without memorizing command-line flags. The browser-based interface provides hardw…

01:08
2026-06-18
byteiota.com
large-language-models

MiniMax M3: Open-Weight Frontier Model at 5% of Opus Cost

MiniMax released the M3 open-weight model, claiming it costs 5% of Claude Opus per task, achieves 59% on SWE-Bench Pro, and supports a 1-million-token context window at one-twentieth the compute of it…

00:00
2026-06-18
techstackups.com
large-language-models

GLM-5.2 vs Claude Opus

Z.ai released GLM-5.2, an open-weights AI model under an MIT license, positioning it between Claude Opus 4.7 and 4.8 in performance while costing less than a fifth of Opus on output tokens. The model …

18:11
2026-06-16
the-ai-corner.com
ai-infrastructure

Inference engineering is the 80% cost cut most teams miss

Inference engineering, the craft of optimizing GPU operations during AI model inference, can cut costs by up to 80% by addressing the split between prefill and decode phases. Two teams using the same …

17:51
2026-06-12
testingcatalog.com
artificial-intelligence

MiniMax M3 launches on NVIDIA platform with Free Endpoint

MiniMax released its M3 multimodal model on NVIDIA's accelerated infrastructure, offering a free public endpoint via NVIDIA's API catalog. The 428-billion-parameter model processes text, images, and v…

16:19
2026-05-29
liquid.ai
large-language-models

Liquid AI reveals 8B-A1B MoE trained on 38T

Liquid AI released LFM2.5-8B-A1B, an edge model designed for fast tool calling on consumer hardware, with a 128K context window and pretraining scaled to 38 trillion tokens. The model, available on Hu…

12:05
2026-05-27
arxiv.org
large-language-models

Stateful Inference for Low-Latency Multi-Agent Tool Calling

Researchers have developed a stateful inference architecture for multi-agent tool calling that reduces per-turn computational cost from full reprocessing to delta-only updates, achieving 2.1x faster p…

← prev page 7 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics