Self-hosting AI does not save money, and I do it anyway Self-hosting the open-weight Qwen3.8 27B model costs $619 per full Artificial Analysis Intelligence Index run through the cheapest zero-data-retention provider, versus $67 for OpenAI's GPT-6 Luna, according to the author's calculations using Artificial Analysis and OpenRouter data as of September 30, 2026. The author reports that Qwen3.8 27B scores 33.7 on the Artificial Analysis Intelligence Index at its xhigh setting, compared with 34.6 for GPT-6 Luna and 31.9 for Claude Opus 4.6, and that running the model at full precision on his own two RTX 3090s costs more in electricity alone than Luna's entire benchmark bill. The author states he self-hosts anyway for fun, sovereignty and privacy, not for savings. I have had this debate many times, so I am finally writing it down. Whenever I say that self-hosting AI does not save money, people hear that I am against self-hosting. That is not the point at all. I am a massive fan of open-weight models. Qwen3.8 27B runs on two RTX 3090s at home https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/ai.nix L43 and 15 more models , and my phone dictation goes to Qwen3-ASR on the same machine https://nijho.lt/post/diction-agent-cli-qwen/ . I think open-source AI is the best thing since sliced bread. I don’t pretend it saves money. I do it for fun, for sovereignty, and for privacy, which I come back to at the end. All benchmark numbers and prices below come from Artificial Analysis https://artificialanalysis.ai/ and OpenRouter https://openrouter.ai/ as of September 30, 2026. I did the math for the hardware I own, and for the counterarguments I hear most. Benchmarks are not everything, and a high score does not reliably predict how a model does on real work. They are still the best we have: all models are benchmaxed, tuned to do well on the popular benchmarks, so the scores at least work as a reference frame for comparing them with each other. My favorite local model right now is Qwen3.8 27B https://artificialanalysis.ai/models/qwen3-8-27b . It came out in August and it is Apache-2.0. People often say models like this run on a single gaming GPU, but that is only true after quantizing them. At full precision, Qwen3.8 27B needs about 54 GB of memory. Quantized to about 4 bits per weight, it fits on one 24 GB card, at some cost in quality. I still split it across both of my 3090s https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/ai.nix L56-L65 , because the second card leaves room for a larger context window. I compare it with OpenAI’s GPT-6 Luna https://artificialanalysis.ai/models/gpt-6-luna , the cheap tier released on September 22. Both models let you choose how long they think, but the settings do not mean the same thing for both. Luna’s token use grows more than 20 times from its lowest setting to its highest, while Qwen’s barely changes. Qwen on low already writes more tokens than Luna on xhigh. So the name of a setting says little on its own. I compare both models at xhigh, the highest setting Qwen offers, and count the tokens each one actually uses. The extra tokens Qwen needs are part of what I am measuring. There, Qwen scores 33.7 on the Artificial Analysis Intelligence Index https://artificialanalysis.ai/methodology/intelligence-benchmarking and Luna scores 34.6. For reference, Claude Opus 4.6, a frontier model from February, scores 31.9. That is an amazing feat in itself. When Opus 4.6 was the best model we had, I never thought that within the same year I could run essentially that level of intelligence in my own house. Artificial Analysis also publishes https://artificialanalysis.ai/leaderboards/models how many tokens each model used to run the index, and what that cost. So the question is simple: what does it cost to run the whole benchmark once? 1 fn:1 Luna runs the whole benchmark for $67. Qwen at full precision, through the cheapest provider https://openrouter.ai/qwen/qwen3.8-27b/providers that does not keep your data ZDR , costs $619. Two things cause that gap: providers charge almost four times as much per output token for Qwen, and Qwen generates almost three times as many tokens to get the same work done. The second bar is my own machine: the electricity alone for running Qwen on my GPUs costs more than Luna’s entire bill. The third bar is what it takes to match the API’s full precision at home: an RTX PRO 6000, a workstation card with 96 GB of memory, enough for the full model.