Member-only story
SGLang started as a UC Berkeley research paper in 2023. By 2026 it was running on 400,000+ GPUs, generating trillions of tokens a day, and had become credible enough that Hugging Face offered it as an alternative once it retired its own inference server. Here’s the actual story.
In December 2025, Hugging Face put its own inference server into maintenance mode. Text Generation Inference, TGI, had been the company’s flagship serving engine since before most people had heard the phrase “LLM inference,” the tool it built to run models on its own Inference Endpoints. It stopped taking new features. Bug fixes only. And Hugging Face’s own documentation started pointing new users somewhere else: to vLLM, or to a tool most people outside machine learning infrastructure have never heard of, called SGLang.
That’s a strange thing for a company to do to its own product. It’s an even stranger thing when the replacement wasn’t built by a competitor with a marketing budget, it came out of a university research lab, published as an academic paper, with a name few people can pronounce with confidence on the first try (it’s “S-G-Lang,” short for Structured Generation Language, not one word). By the time Hugging Face quietly retired TGI, that research project was running on more than…