Local LLM Inference at Scale with vLLM A developer evaluated vLLM, an open-source serving engine for self-hosted open-weight models, benchmarking continuous batching, prefix caching and structured outputs on an NVIDIA DGX Spark with a GB10 Grace Blackwell chip and 128 GB of unified memory. The writeup explains how PagedAttention stores the KV cache in fixed-size blocks to avoid fragmentation and how continuous batching keeps GPU capacity busy, and it presents a throughput table for recent open-weight models under 120B parameters. This post explore vLLM, a powerful, full featured serving engine for your open weight models. Checkout our previous post on llama.cpp and Ollama https://data4sci.substack.com/p/self-hosting-llms-with-llamacpp for a fuller picture of how to self-host models. And of course, don’t forget to and to help us grow A while ago we explored self-hosting LLMs with llama.cpp and Ollama https://data4sci.substack.com/p/self-hosting-llms-with-llamacpp , a family of tools that make running a local model just a matter of a few commands. This post continues this line of exploration with vLLM https://github.com/vllm-project/vllm , a full featured self-hosting engine that is capable of scaling to thousands of requests. We will build a quantitative picture by evaluating the vLLM features that matter in practice: continuous batching, prefix caching and structured outputs by using the OpenAlex corpus of scaling-laws papers we first explored in our MCP server post https://data4sci.substack.com/ . Throughput math LLMs work by generating text token by token. Generating token t+1 requires a forward pass across the entire set of weights with the first t tokens as context. In other words, a 7-billion-parameter model in 16-bit precision is 14 GB of weights see the Ollama https://data4sci.substack.com/p/self-hosting-llms-with-llamacpp post for the calculation . The bottleneck is the memory-bandwidth, so a quick back of the envelope calculation of generation speed is: Fortunately, in most models the weights are the same for every request. We can in principle processing N similar requests in a single forward pass while still only needing to reading the weights once. Batching gives us free performance improvements up to the point where the hardware simply can’t handle the calculations fast enough. A server handling many users is significantly more efficient per token than a notebook generating one answer at a time. Attention is… what makes this rosy picture not so rosy. Transformers use a KV cache https://huggingface.co/blog/not-lain/kv-caching to avoid recomputing attention over the prefix at every step. The KV cache grows with as the sequence grows, since for each request it keeps the keys and values of every previous token resident in GPU memory. Naive implementations pre-allocate the maximum size necessary, causing batch sizes to be limited to a handful of requests. vLLM introduced two foundational ideas to improve on these limitations: - PagedAttention takes a page from virtual-memory pages