Bringing AI Home: Why Regulated Enterprises Are Moving from Cloud GPUs to Local Blackwell Machines A regulated enterprise moved its AI inference workloads off rented cloud GPUs onto a desktop-sized NVIDIA GB10 Grace Blackwell machine, sold as the DGX Spark and in partner versions such as the ASUS Ascent GX10, citing unpredictable usage-shaped GPU costs, network latency, and compliance audit exposure across four network hops, two vendors and two jurisdictions. The team runs open-weight models on the local box with vLLM for production and Ollama for developer machines, both exposing an OpenAI-compatible HTTP API, and still rents cloud GPUs for one-off training runs and large benchmarks. The author argues most regulated enterprises in finance and healthcare will run AI this way within five years, treating the machine as a departmental inference appliance rather than a cheap H100. We moved our AI workloads off rented cloud GPUs and onto a desktop-sized Grace Blackwell machine. Here is what we learned, what it cost, what broke, and why I think most regulated enterprises will run AI this way within five years. It started with two questions nobody on the team could answer cleanly. The first came from finance: why does our GPU bill look like a second payroll? We were paying for managed AI services in one hyperscaler and spinning up on-demand GPU pods on a second provider for experiments. Each was reasonable on its own. Together they were an unpredictable, usage-shaped cost that grew every time an engineer left an endpoint warm over the weekend. The second came from a compliance reviewer, and it was harder: where, exactly, does the data go between the microphone and the model? We could answer it. But the honest answer had four network hops, two vendors, two jurisdictions and a stack of agreements in it. In a regulated industry, every one of those is a sentence someone has to defend in an audit. So we tried something that would have sounded eccentric three years ago. We moved inference onto a machine that fits on a desk. This article is the playbook I wish I had read first. It is written for engineers, architects and technology leaders in finance and healthcare who are weighing the same move. No vendor pitch, and nothing proprietary: just the architecture, the traps, and a blueprint you can copy. No single factor justified the move. Four smaller ones, pushing in the same direction, did. 1. Cost became a shape, not a number. Cloud GPU pricing is per hour or per token. That is perfect for bursty experiments and painful for steady, predictable inference. Once a workload runs every working day, a one-time hardware purchase starts to win, and it keeps winning every month after it pays for itself. 2. Latency is physics. Every round trip to a remote region adds delay you cannot optimise away. For interactive AI, where a person is waiting on the answer, a model one switch away feels different from one behind the public internet. 3. Compliance is a diagram, and shorter diagrams win. Health data HIPAA, GDPR, national health-data laws and financial data banking secrecy, data-residency rules, the EU AI Act, DORA all ask the same question: who can see this, and where does it live? “It never leaves the building” is the shortest answer an auditor can hear. 4. Control over the model itself. Hosted models change under you. A provider retires a version or adjusts behaviour, and your carefully measured accuracy is no longer the accuracy you ship. With open-weight models on your own hardware, the model you validated is the model you run until you decide otherwise. None of this makes the cloud wrong. It makes it the wrong default for steady, sensitive inference. We still use rented GPUs for one-off training runs and large benchmarks. The difference is that they are now a tool we reach for, not a place we live. The machine class that made this possible is NVIDIA’s GB10 Grace Blackwell desktop platform, sold as the DGX Spark and in partner versions such as the ASUS Ascent GX10. The headline is not raw speed. It is memory . Where it falls short, honestly: The right mental model: not a cheap H100, but a departmental inference appliance . Big enough to host a small portfolio of specialised models for one team, product or clinic, and small enough to sit inside your own network boundary. The stack that works is deliberately unexciting. Each layer is open source, replaceable, and speaks a standard protocol. Serving: vLLM for production, Ollama for the laptop. Ollama is the fastest way to try a model on a developer machine. vLLM is what you want on the shared box: continuous batching, paged attention, quantised weights, and an OpenAI-compatible API. Both speak the same HTTP dialect, so the application code does not care which one answers. One model per port. Run each model as its own vLLM process on its own port, with its own memory share. It looks wasteful and is the opposite: one model crashing or being upgraded never takes the others down, and you can see exactly which model is using what. Schema-constrained output. If your application expects JSON, pass the JSON schema to the server vLLM’s guided decoding, Ollama’s format . The model then cannot return malformed output. In our measurements this changed no answers and cost no speed, and removed a whole class of parsing failures. Models: open weights, chosen per job. Instead of one giant model doing everything, we run a small portfolio: a general instruction model for structured extraction, a separate model for translation, a speech-recognition model, and a domain-tuned model under evaluation. Families worth evaluating today include Qwen, Mistral, Gemma including the medical MedGemma variants , Llama and Whisper for speech. Smaller specialists beat one generalist on both speed and accuracy more often than you would expect. Quantisation: do it yourself if you must. Not every model ships in a 4-bit format your server can read. Tools like LLM Compressor can produce one in about an hour on the box itself. Then measure the quantised model against the full-precision one, because “4-bit is basically the same” is a hypothesis, not a fact. Operations: systemd, not Kubernetes. For one or two boxes, a small set of systemd units and a mode-switch script beat an orchestrator. Kubernetes earns its keep at ten machines, not two. 1. Measure, never assume. Every model swap, quantisation and prompt change gets scored against a labelled evaluation set before it ships. Several “obvious upgrades” a bigger model, a newer model, a cleaner audio pipeline measured worse on our data. Without the eval set we would have shipped all of them. 2. Bigger is not better; fitter is better. A mid-sized model that is fast enough for a person to wait on beats a giant one that is right a little more often but arrives too late to be used. Latency is an accuracy problem when a human is in the loop. 3. Budget memory like money. Write down each model server’s share of memory, add the per-process overhead, and leave headroom for the operating system. Our first plan summed to under 90% and still left the box gasping, because each server also took host memory beyond its share. 4. Fail closed, and degrade gracefully. Decide in advance what happens when a model is down. For an advisory feature: say “temporarily unavailable”, never quietly swap in a different model nobody validated. For a core workflow: fall back to a simpler, rule-based path so work continues coarsely rather than stopping. 5. Watch the network you didn’t think about. “Local” only means local if the box is on your premises. If you reach your GPU server over the internet, your compliance story changes, even if you own the hardware. Write down which hosts count as on-premises, and have the system report it. 6. Record provenance on every output. Store which model, which version, which prompt and which configuration produced every AI artifact, captured when the job starts . Months later, when someone asks why a document says what it says, “whatever was configured at the time” is not an answer. 7. Benchmarks are not your data. Public leaderboards are a shortlist, not a decision. A model that leads a medical benchmark can still loop, hedge or invent on your real inputs. Your own labelled cases, reviewed by domain experts, are the only scoreboard that counts. Everything that touches sensitive data sits inside the boundary. The cloud becomes an opt-in tool for work that carries no personal data. 1. Write the evaluation set before you buy anything. Collect 50–200 real, labelled examples of the task, reviewed by a domain expert. Define the metric that decides “good enough”. This set will choose your model, your quantisation and your hardware for you. 2. Size memory with arithmetic, not hope. Weights take roughly the parameter count times the bits per weight divided by eight, before the KV cache and server overhead: weights GB ≈ parameters billions × bits per weight ÷ 8 A 70B model at 4-bit is about 35 GB of weights; an 8B model at 16-bit is about 16 GB. Add 20–30% per server for cache and overhead, and keep 10–15% of the box free. 3. Place the box on premises, on its own network segment. No inbound internet. Administrators reach it over VPN. Only the application backend may call the model ports. 4. Pin the base layer. Use the vendor’s OS image or a supported Ubuntu, with the NVIDIA driver, CUDA and container toolkit at versions you write down and upgrade deliberately. 5. Serve each model as its own process, behind a key. vllm serve