cd/entity/SGLang· home› entities› SGLang
grep -l @sglang /news/*.json | wc -l → 236

SGLang

mentions 236 type Organization page 1/12 feed RSS

// recent coverage 236 mentions

00:00
2026-10-03
mindstudio.ai
large-language-models

IQuest-Q1: How to Self-Host the 320B Agentic Coding Model

IQuest released IQuest-Q1, an open-weight 320B-parameter Mixture-of-Experts coding model that activates only 15B parameters per token and supports a 524,288-token context window, with official deploym…

06:11
2026-10-02
byteiota.com
ai-tools

Janus: Run GGUF Models Locally With One Go Binary

Janus, an MIT-licensed single Go binary from the Vibra-Ingenn project, launched as a Show HN on October 1 with 42+ points, wrapping llama.cpp's Vulkan backend to run GGUF models on AMD, Intel, or NVID…

14:51
2026-10-01
arxiv.org
artificial-intelligence

Context Language Models

A paper submitted to arXiv on 29 Sep 2026 introduces Context Language Models (CLMs), language models that natively manage their own context by treating the context as a file the model can update witho…

00:00
2026-09-30
mindstudio.ai
large-language-models

IQuest-Q1: Inside the 320B MoE Model Built for Agentic Coding

IQuest released IQuest-Q1, an open-weight 320B-parameter Mixture-of-Experts language model with 15B active parameters per token and a 512K-token context window, built for agentic coding and multi-step…

00:00
2026-09-30
mindstudio.ai
large-language-models

How to Deploy IQuest-Q1 with SGLang or vLLM

IQuest released IQuest-Q1, a 320-billion-parameter Mixture-of-Experts model with roughly 15 billion active parameters per token, an 88-layer transformer, 256 experts (8 active), and a 524,288-token co…

07:00
2026-09-29
haoailab.com
ai-infrastructure

UniServe: Serving FastH3 at Its Fastest

UniServe, a serving engine for FastH3 8-Step text-to-video-with-audio generation from the hao-ai-lab, delivers lower median end-to-end latency and 20–46% higher throughput than FastVideo, vLLM-Omni an…

00:00
2026-09-28
int21.ai
ai-infrastructure

An AlphaGo Moment for Inference?

INT21 generated 20 inference engines across seven model categories in two weeks using Rust with C++ and CUDA components, and its MiMo engine reached 1,308 tokens/s versus 540 for tuned SGLang and 1,01…

22:01
2026-09-26
pub.towardsai.net
ai-safety

[Framework] Zero Standing Privilege for AI Workloads

Enterprise data breaches involving unsanctioned Shadow AI cost organizations an average of $650,000 more than conventional security incidents, and 20.0% of global enterprises have already suffered a p…

00:00
2026-09-26
mindstudio.ai
artificial-intelligence

MiMo-V2.6-Flash-RL: Xiaomi's Efficient 309B Omnimodal Model

Xiaomi released MiMo-V2.6-Flash-RL, a sparse Mixture-of-Experts omnimodal model with 309 billion total parameters and 15 billion active per token, positioned as the efficiency-focused sibling to MiMo-…

00:00
2026-09-25
mindstudio.ai
artificial-intelligence

MiMo-V2.6-Pro-RL: Xiaomi's 1T-Parameter Agentic Model, Explained

Xiaomi's MiMo team released MiMo-V2.6-Pro-RL, an open-weight trillion-parameter mixture-of-experts agentic model with 1.02 trillion total parameters, 42 billion active per token, and a 1 million token…

14:08
2026-09-24
huggingface.co
large-language-models

Accelerating vision-language models with LFM2.5-VL-DSpark

LiquidAI released LFM2.5-VL-DSpark, a 279.5M-parameter speculative decoding drafter for its LFM2.5-VL-3B vision-language model that adds 8.9% to the target model's parameter count while delivering dec…

16:34
2026-09-23
nunchux.ai
generative-ai

Nunchux on AMD MI355X: 5s MiniMax-H3 Videos in 1.3s

Nunchux generated a 5-second MiniMax-H3 video in 1.33 seconds on a single server with eight AMD MI355X GPUs, a 21.8× speedup over SGLang on the same hardware, according to the company's benchmark post…

00:00
2026-09-23
doug.sh
ai-tools

Custom Models in Oh My Pi: vLLM, llama.cpp, SGLang and More

Oh My Pi (omp) 18.2.7 changed how custom local models are configured, requiring users to rename the provider from `local` to an unused name and add `qwenTemplateReasoningEffort: true` to a model's `co…

page 1 / 12 next →
// co-occurs with top 8 entities
// topics top 6 topics