cd/entity/SGLang· home entities SGLang
grep -l @sglang /news/*.json | wc -l → 145

SGLang

mentions 145 type Organization page 4/8 feed RSS

// recent coverage 145 mentions

12:08
2026-07-27
sourcefeed.dev
artificial-intelligence

The Real Cost of 'Just Use vLLM,' According to Netflix

Netflix's AI Platform team published a detailed account of its LLM serving platform, revealing that version pinning between NVIDIA Triton Inference Server and vLLM, a Python GIL bottleneck, and KV-cac…

09:08
2026-07-24
sourcefeed.dev
artificial-intelligence

KTransformers Turned CPU Offloading Into Real Infrastructure

KTransformers, a Tsinghua-born framework now with 19k GitHub stars, has restructured from a standalone inference system into CPU kernels and placement strategies embedded within SGLang and LLaMA-Facto…

09:03
2026-07-23
ktransformers.net
artificial-intelligence

KTransformers – Flexible LLM Inference Framework

KTransformers, a flexible LLM inference framework, enables deployment of 100B+ parameter models locally on a single RTX 5090 (32GB VRAM) using CPU/GPU heterogeneous computing without quantization. The…

12:01
2026-07-22
pub.towardsai.net
large-language-models

The Complete Technical Guide to Running LLMs Locally in 2026

A technical guide to running large language models locally in 2026 provides hardware math, quantization tradeoffs, and benchmarks of five inference engines, with case studies from the author's 16GB Ap…

03:55
2026-07-19
henrypan.com
ai-agents

Harness Training

A developer known as workofart has created a PyTorch-like framework for training AI agent harnesses, achieving a 45-minute experiment cycle on terminal bench tasks by enforcing deterministic inference…

00:51
2026-07-17
tinfoil.sh
ai-infrastructure

Architecting Secure Prompt Caching

Tinfoil announces cached prompt pricing in its Inference API, a feature that reduces compute for eligible requests by caching recently processed inputs, but the company warns that caching introduces t…

21:12
2026-07-15
thinkingmachines.ai
artificial-intelligence

Inkling Model Card

Thinking Machines Lab, Inc. released Inkling, a general-purpose multimodal model with 975 billion total parameters and 41 billion active parameters, on July 15, 2026 under an Apache 2.0 license. The m…

00:00
2026-07-15
dibi8.com
artificial-intelligence

SGLang — Structured Generation and Fast LLM Serving Engine

SGLang, an open-source LLM inference engine, introduces RadixAttention for prefix caching and grammar-constrained decoding, achieving 25x throughput improvement over vLLM for structured output tasks. …

00:00
2026-07-15
huggingface.co
large-language-models

Welcome Inkling by Thinking Machines

Thinking Machines released Inkling, a ~1T-parameter multimodal Mixture-of-Experts LLM with 1M context window and native support for image, audio, and text inputs, now available on Hugging Face. The mo…

← prev page 4 / 8 next →
// co-occurs with top 8 entities
// topics top 6 topics