I run a 27B local model on my own desktop. Not a server, not a headless box: the same Windows 11 machine I take calls on and write on all day. It has an RTX 5090 with 32 GB of VRAM, vLLM running in WSL2, and it serves Qwen3.8 27B, NVFP4-quantized, to a DeepSeek Harness (DSH) agent that runs multi-hour coding sessions: shell, edits, tests, subagents, web search, and document work.
This machine was one of three in a separate comparison I wrote up: DGX Spark, RTX 5090, Radeon R9700. The 5090 was the fastest of the three in long coding sessions. It was also the one that collapsed to single-digit tokens per second during each of three video calls.
The part worth writing up is that nothing failed. vLLM never threw an out-of-memory error and never crashed. When the desktop needed VRAM, generation slowed to a fraction of its normal rate, and the token rate was the only sign. If you expect a GPU that runs out of memory to tell you so, this is the case where it didn’t.
This is the diagnosis, end-to-end: the symptoms, the test that isolated the cause, the VRAM math, and the configuration change that stopped it. Every number is from that one machine and my workload. Where the evidence doesn’t fully close a question, I say so.
The diagnostic that pointed at memory, separate from the fix: a short-request run stuck at 10–11 tok/s, with KV cache usage at 1.7%, returned to 82.8 tok/s after restarting only vLLM.
The tok/s figures are observed DSH session rates and vLLM reporting intervals on this one machine, not a controlled benchmark. Below: how I got from “it’s slow” to that table.
The pattern was what first pointed me toward memory pressure rather than the model itself.
One DSH session began around 77 tokens per second. By the end, the session average was around 27, and that average includes the stretch at 77, so the late rate was far lower. A subagent running on the same server during the collapse measured around 7. The agent never hung. It kept producing work, just slowly.
That session overlapped two Teams calls, and each call brought a slowdown. Twice during Teams calls while the agent was running, the screen blacked out or flickered, Teams crashed, and the video failed on rejoining. The agent and the desktop were fighting over the same card. In troubleshooting, the card showed 32,049 MiB in use, nearly its entire 32 GB.
The telemetry did not indicate thermal or power throttling. The card was not overheating. It was full.
After the second call caused the second slowdown, a short-request run had fallen to roughly 10–11 tokens per second. During that slow run, vLLM reported KV cache usage at just 1.7%.
That number rules out KV-cache exhaustion. And because this was a short-request run, not a huge prefill, prompt length wasn’t a credible explanation either.
Then I restarted only vLLM. Not the desktop. Not the model. Not the configuration. One process. The same run returned to 82.8 tokens per second.
The restart doesn’t say which serving state was the problem. A fresh process reloads the weights and reallocates the cache, so it would fix either one. What it does say is that the model, the prompt, and the client were fine, and the serving process’s memory state was not.
Then I took a third call with the configuration unchanged. Performance took a nose dive again.
Call, collapse. Restart, recovery. Call, collapse. That sequence is the diagnosis: the trigger was the desktop’s GPU demand during calls, and what suffered was vLLM’s memory state.
The flag at the center of it is — — gpu-memory-utilization. It sets vLLM’s GPU memory budget for this instance: the share of the card that vLLM can treat as its own. I had it at 0.93. On a 32 GB card, that leaves a little over 2 GB for everything else the card does: the Windows desktop, the display compositor, and Teams’ video encoding.
That split is fine on a quiet desktop. It stops being fine the moment a video call actually uses the GPU.
When the call added GPU demand, my working explanation is that some of vLLM’s allocations lost physical VRAM residency: pushed out of VRAM to make room for the desktop. Under WSL/Windows GPU memory management, that can mean data is served from system memory instead, introducing PCIe traffic into the decode path, which would explain a slowdown of this size.
Stated as a ladder, because the levels matter:
-
Observed: the card is nearly full; a desktop workload is triggering the slowdown; a restart restores performance; low KV utilization rules out cache saturation; more headroom prevents the severe recurrence.
-
Strong inference: contention for physical VRAM is the failure mode.
-
Hypothesis: some serving allocation was evicted or migrated into host-visible memory, and the resulting PCIe traffic caused the token-rate loss.
I did not trace any migration into system memory, and I did not isolate which allocations lost residency (weights, KV cache, or both). The PCIe detail is the rung that best fits the observations, not something I measured. I’m not going to pretend it is.
“The model fits” was never the issue; the model always fit. The question is how much VRAM remains after the model loads, because a model that fits in VRAM is not the same as a serving configuration with sufficient headroom. A 27B model at two bytes per parameter needs about 54 GB for weights alone in FP16/BF16, so on a 32 GB card, it has to be quantized.
Two NVFP4 builds, measured in a 16,384-token judge setup:
A note on the columns, because they look like they don’t add up: “Left for KV cache” is a hand calc, budget minus weights. vLLM doesn’t hand all of that to the cache; it first carves out a fixed chunk for activation and graph memory (roughly 1.5 GiB at these settings, as the two rows imply), and that fixed cost eats a larger share of the small Unsloth pool than of the ModelOpt one. The bytes per cached token are the same in both builds; what differs is how much of the leftover each pool actually gets.
Those are configuration-specific cache capacities, not guaranteed context lengths. The ModelOpt build’s smaller weight footprint is why my long agent sessions have meaningful KV headroom at all, and the FP8 KV cache in my launch script halves the bytes per cached token on top of that.
The issue was never the model. It was who keeps residency when the desktop competes for the same 32 GB, and at a 0.93 budget, there wasn’t enough headroom for the desktop and the serving process to coexist cleanly.
Two changes were made after the third call.
-
— — gpu-memory-utilization: 0.93 → 0.85. This is the intervention aimed at desktop contention: give the card room to breathe for Windows and video calls, returning roughly 2.5 GiB to the desktop. -
— — max-model-len: 262,144 → 163,840. This is a workload guardrail, not a memory fix: it caps the length of a single request. To be clear about what it does not do: it does not give VRAM back to the desktop. The memory budget is the flag that does that.
Both are precautions against the diagnosed memory-pressure failure mode. I changed them together, so I can’t attribute the result to either one alone.
The fix has a price. The lower budget takes about 2.5 GiB away from vLLM, more than a quarter of the 9.26 GiB KV pool this build had in the judge setup, and the lower ceiling gives up about 98,000 tokens of maximum request length. For my workload (sustained agent sessions that compact as they run), that trade was worth it.
Since making both changes, I have taken multiple Teams calls with the agent running and have not recreated the severe collapse. The worst drop I have seen since is roughly 83 to 52 tokens per second. Fair warning: the 5090 still has more throughput volatility in my use than the Spark or the Radeon.
Two things I would check.
The launch script, sanitized:
#!/usr/bin/env bashset -euo pipefail# Initialize Conda for non-interactive shellsource /path/to/miniconda3/etc/profile.d/conda.shconda activate env_vllmexport CUDA_HOME=/path/to/miniconda3/envs/env_vllm/lib/python3.12/site-packages/nvidia/cu13export PATH="$CUDA_HOME/bin:$PATH"export LIBRARY_PATH="/usr/lib/wsl/lib:${LIBRARY_PATH:-}"export VLLM_USE_V2_MODEL_RUNNER=0vllm serve \ gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 \ --host 0.0.0.0 \ --port 8000 \ --served-model-name qwen3.8-27b \ --seed 0 \ --tensor-parallel-size 1 \ --trust-remote-code \ --max-model-len 163840 \ --kv-cache-dtype fp8 \ --gpu-memory-utilization 0.85 \ --max-num-seqs 2 \ --max-num-batched-tokens 8192 \ --enable-prefix-caching \ --enable-chunked-prefill \ --reasoning-parser qwen3 \ --default-chat-template-kwargs '{"preserve_thinking":false,"reasoning_effort":"medium"}' \ --enable-auto-tool-choice \ --tool-call-parser qwen3_xml \ --generation-config vllm \ --override-generation-config '{ "temperature": 0.8, "top_k": 40, "top_p": 0.95, "min_p": 0.05, "repetition_penalty": 1.0 }'
Two flags to change before you copy this. — — host 0.0.0.0 listens on every network interface with no API key, so bind to 127.0.0.1 (my agent connects over loopback) or add — — api-key if anything else on your network can reach the machine. And — — trust-remote-code runs code shipped with the checkpoint, so know what you’re trusting.
env_vllm is a pinned conda environment (Python 3.12.14, vLLM 0.29.0, torch 2.13.0+cu130), rebuilt from an environment file and a requirements lock exported from the working machine. The CUDA_HOME and PATH exports put that environment’s own CUDA toolkit on the load path, LIBRARY_PATH is the WSL piece, and VLLM_USE_V2_MODEL_RUNNER=0 pins the non-v2 model runner.
The checkpoint is the 5090-oriented NVFP4 build. Its model card notes that the MTP head was removed, so the script doesn’t enable speculative decoding. If you serve a different build, such as the Unsloth NVFP4 checkpoint, re-run the headroom math before you trust my numbers.
A 32 GB card shared with a desktop is a negotiation. At a 0.93 vLLM budget, mine crossed the line three times without vLLM ever reporting an error. Dropping the budget to 0.85 gave Windows enough headroom that I haven’t reproduced the severe collapse since.
The caveats, because they’re the difference between a field report and a press release: one machine, one workload, my measurements, and no isolation of which serving state left the card.
This post is the VRAM diagnosis from my three-platform comparison, One Harness, One Agent, Three GPUs: Running Qwen on a DGX Spark, an RTX 5090, and a Radeon R9700. In that comparison, the 5090 was the fastest of the three in long coding sessions. The fix above is what kept that speed from turning into a black screen mid-call.
My RTX 5090 Fell from 77 to 7 Tokens Per Second. VRAM Pressure Was the Culprit was originally published in Towards AI on Medium, where people are continuing the conversation by highlighting and responding to this story.