cd/entity/SGLang· home› entities› SGLang
grep -l @sglang /news/*.json | wc -l → 236

SGLang

mentions 236 type Organization page 3/12 feed RSS

// recent coverage 236 mentions

04:11
2026-09-12
skydiscover-ai.github.io
ai-agents

Building Specialized Systems We Can Trust with Agents

SkyDiscover-Synthesize (SkySynth), an autonomous pipeline that builds just-in-time specialized systems, synthesizes key-value stores up to 2.3× faster than Redis and FASTER, inference engines with up …

22:32
2026-09-10
skeptrune.com
ai-infrastructure

The Inference Engineering Skills Map

A former B2B SaaS web developer who switched to inference engineering about a month ago published a skills map arguing the field requires only three service categories: routing and scheduling via Nvid…

13:39
2026-09-10
github.com
ai-infrastructure

LRU is harder to beat than the KV-cache papers suggest

A prefix-cache simulator replaying 68,266 requests from 393 real Claude Code sessions and 23,608 Mooncake requests failed to beat the production LRU baseline in three separate attempts, according to t…

00:00
2026-09-10
mindstudio.ai
ai-products

Nex-N2.5 Mini Hands-On: Testing Next AGI's Agentic Model

Next AGI released Nex-N2.5 Mini, a multimodal agentic model built on a Qwen3.5 mixture-of-experts backbone, under an Apache 2.0 license on Hugging Face, requiring two 80GB-class GPUs to run. In hands-…

00:00
2026-09-10
mindstudio.ai
ai-products

How to Run Nex-N2.5 Mini Locally on RunPod (Dual H100 Setup)

Nex AGI's Nex-N2.5 Mini, the smaller model in the company's N2.5 agentic family, requires two 80GB-class GPUs such as H100s and consumed roughly 66GB of VRAM per card in testing on RunPod, according t…

12:01
2026-09-09
pub.towardsai.net
large-language-models

Why Is Your LLM Recomputing the Same Prompt 1,000 Times a Day?

Production LLM traffic is dominated by repeated prompt prefixes—system prompts, chat history, and shared documents—causing inference engines to recompute identical KV cache state thousands of times da…

05:56
2026-09-09
forum.level1techs.com
artificial-intelligence

Local LLM Based Coding on HPC Equipment

A developer running custom simulation software has moved from cloud-based AI coding assistance to local inference on high-end HPC equipment, using Hermes, Ollama, vLLM, and SGLang with models like Qwe…

02:11
2026-09-09
promptcube3.com
machine-learning

Why India needs more ML infrastructure builders and fewer API

At a recent event, speakers including Aritra Roy Gosthipaty from Hugging Face emphasized the need for India to focus on machine learning infrastructure builders rather than API-level developers, highl…

22:09
2026-09-08
frontierroles.com
artificial-intelligence

Staff Applied AI Inference Engineer — Crusoe

Crusoe, a vertically integrated AI infrastructure company, is hiring a Staff Applied AI Inference Engineer in San Francisco with a salary range of $215,000–260,000 per year, which sits 8% above the $2…

00:00
2026-09-08
rocm.blogs.amd.com
artificial-intelligence

veRL on AMD: Production-Ready RL Post-Training on ROCm

AMD and the veRL project released a production-ready reinforcement learning post-training container for AMD Instinct GPUs, supporting MI300 and MI355 series on ROCm, with AITER-accelerated vLLM and SG…

00:00
2026-09-08
mindstudio.ai
artificial-intelligence

How to Run MiniCPM5-2B Locally with SGLang or llama.cpp

ModelBench's MiniCPM5-2B, a 2-billion-parameter dense language model from the Tsinghua University spinout, can run locally with SGLang or llama.cpp, requiring roughly 4GB of VRAM for weights but up to…

13:01
2026-09-07
pub.towardsai.net
ai-infrastructure

Superlinked Inference Engine

Superlinked has released the Superlinked Inference Engine (SIE), an open-source server designed to consolidate the many small models used in AI agent workflows into a single deployment, addressing wha…

05:00
2026-09-07
marktechpost.com
artificial-intelligence

IFM Releases K2 Horizon: Six Apache 2.0 Models From 0.9B to 375B

The Institute of Foundation Models (IFM), the frontier lab launched by MBZUAI in May 2025, released K2 Horizon, a fleet of six Apache 2.0-licensed open-source models ranging from 0.9B to 375B paramete…

← prev page 3 / 12 next →
// co-occurs with top 8 entities
// topics top 6 topics