{"slug": "self-hosting-ai-does-not-save-money-and-i-do-it-anyway", "title": "Self-hosting AI does not save money, and I do it anyway", "summary": "Self-hosting the open-weight Qwen3.8 27B model costs $619 per full Artificial Analysis Intelligence Index run through the cheapest zero-data-retention provider, versus $67 for OpenAI's GPT-6 Luna, according to the author's calculations using Artificial Analysis and OpenRouter data as of September 30, 2026. The author reports that Qwen3.8 27B scores 33.7 on the Artificial Analysis Intelligence Index at its xhigh setting, compared with 34.6 for GPT-6 Luna and 31.9 for Claude Opus 4.6, and that running the model at full precision on his own two RTX 3090s costs more in electricity alone than Luna's entire benchmark bill. The author states he self-hosts anyway for fun, sovereignty and privacy, not for savings.", "body_md": "I have had this debate many times, so I am finally writing it down.\n\nWhenever I say that self-hosting AI does not save money, people hear that I am against self-hosting.\nThat is not the point at all.\nI am a massive fan of open-weight models.\n[Qwen3.8 27B runs on two RTX 3090s at home](https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/ai.nix#L43) (and 15 more models), and my phone dictation goes to [Qwen3-ASR on the same machine](https://nijho.lt/post/diction-agent-cli-qwen/).\nI think open-source AI is the best thing since sliced bread.\n\nI don’t pretend it saves money. I do it for fun, for sovereignty, and for privacy, which I come back to at the end.\n\nAll benchmark numbers and prices below come from [Artificial Analysis](https://artificialanalysis.ai/) and [OpenRouter](https://openrouter.ai/) as of September 30, 2026.\nI did the math for the hardware I own, and for the counterarguments I hear most.\n\nBenchmarks are not everything, and a high score does not reliably predict how a model does on real work. They are still the best we have: all models are benchmaxed, tuned to do well on the popular benchmarks, so the scores at least work as a reference frame for comparing them with each other.\n\nMy favorite local model right now is [Qwen3.8 27B](https://artificialanalysis.ai/models/qwen3-8-27b).\nIt came out in August and it is Apache-2.0.\nPeople often say models like this run on a single gaming GPU, but that is only true after quantizing them.\nAt full precision, Qwen3.8 27B needs about 54 GB of memory.\nQuantized to about 4 bits per weight, it fits on one 24 GB card, at some cost in quality.\nI still [split it across both of my 3090s](https://github.com/basnijholt/dotfiles/blob/e63a3f341ff36b7b57bf31361c2844e1d8b78e95/configs/nixos/hosts/pc/ai.nix#L56-L65), because the second card leaves room for a larger context window.\n\nI compare it with OpenAI’s [GPT-6 Luna](https://artificialanalysis.ai/models/gpt-6-luna), the cheap tier released on September 22.\nBoth models let you choose how long they think, but the settings do not mean the same thing for both.\nLuna’s token use grows more than 20 times from its lowest setting to its highest, while Qwen’s barely changes.\nQwen on low already writes more tokens than Luna on xhigh.\n\nSo the name of a setting says little on its own.\nI compare both models at xhigh, the highest setting Qwen offers, and count the tokens each one actually uses.\nThe extra tokens Qwen needs are part of what I am measuring.\nThere, Qwen scores 33.7 on the [Artificial Analysis Intelligence Index](https://artificialanalysis.ai/methodology/intelligence-benchmarking) and Luna scores 34.6.\nFor reference, Claude Opus 4.6, a frontier model from February, scores 31.9.\nThat is an amazing feat in itself.\nWhen Opus 4.6 was the best model we had, I never thought that within the same year I could run essentially that level of intelligence in my own house.\n\nArtificial Analysis also [publishes](https://artificialanalysis.ai/leaderboards/models) how many tokens each model used to run the index, and what that cost.\nSo the question is simple: what does it cost to run the whole benchmark once?[1](#fn:1)\n\nLuna runs the whole benchmark for $67.\nQwen at full precision, through the [cheapest provider](https://openrouter.ai/qwen/qwen3.8-27b/providers) that does not keep your data (ZDR), costs $619.\nTwo things cause that gap: providers charge almost four times as much per output token for Qwen, and Qwen generates almost three times as many tokens to get the same work done.\n\nThe second bar is my own machine: the electricity alone for running Qwen on my GPUs costs more than Luna’s entire bill.\n\nThe third bar is what it takes to match the API’s full precision at home: an RTX PRO 6000, a workstation card with 96 GB of memory, enough for the full model.<sup>[2](#fn:2)</sup>\nIts electricity alone costs twice Luna’s bill, and one run takes more than seven weeks.\nThe card by itself sells for $10,000 to $20,000, so with a computer around it, you are paying for a pretty nice car.\nBuy a few of them to run agents in parallel, and you are building a small data center of your own.\n\nThe bars for hardware at home are electricity only, so they are the lowest these runs can cost: the cards come on top, and how much depends on how busy you keep them, which is what section 4 is about.\n\nWhat I care about is the cheapest way to get answers of Qwen’s quality, from whichever model gives them.\n\nThis is the argument I hear most, so assume the hardware is free and only the electricity counts.\n\nI measured my own machine.\nIt runs Qwen in vLLM, split across both cards, at about 107 tokens per second, while the machine draws about 700 W.<sup>[3](#fn:3)</sup>\nIt then needs four weeks, running day and night, to get through the benchmark once.\nI also give my quantized copy the score Artificial Analysis measured through Alibaba’s API.\nI have not rerun the benchmark on it, and anything lost to quantization makes my machine look worse.\n\nAbove 14.5 cents per kWh, my electricity alone costs more than Luna’s API bill. Below it, I save a few dollars per run and wait four weeks instead of a few hours.\n\nThose four weeks matter more than the few dollars. An API takes hundreds of requests at once, so even one benchmark run can finish in an hour or two. My setup works on one conversation at a time: when I sent it eight requests at once, seven of them waited in line.\n\nThis is one more reason my coding agents run on APIs.\nWhen I [run several agents in parallel](https://nijho.lt/post/parallel-agentic-coding/), each of them should be as fast as if it were alone.\n\nWhen my 3090s generate a token, they read all of the model’s weights from memory to produce one token for one conversation. Most of the chip’s compute sits idle while it waits for memory. A datacenter GPU reads the same weights once and produces a token for hundreds of conversations at the same time. This is called batching, and it is where most of the efficiency comes from.\n\nArtificial Analysis measures this with [AgentPerf](https://artificialanalysis.ai/hardware-inference-stack/datacenter), which replays real coding-agent sessions against datacenter hardware.\nIt counts how many agents a system can serve while each one still gets at least 60 tokens per second.\n\nThe rack of 36 GB300s serves 20 times more agents per kW than the 8-GPU H200 server. Part of that is newer chips, and part is software: NVIDIA tuned the B300 and GB300 setups itself, while Artificial Analysis configured the H200 one. The B300 and the GB300 share both the chip generation and the software. Most of the gap between those two bars comes from the rack itself: 36 GPUs on one fast interconnect keep more than a thousand agents in flight at once. My two 3090s serve one conversation at a time.\n\nSolar power is not free. With net metering, every kWh my GPUs burn is a kWh that does not get credited at the retail price. Without net metering, it is a kWh I do not sell, unless the surplus would go to waste anyway, which is the free-power case below. And the GPUs also run at night.\n\nBut fine, say the electricity is free and the only cost is the hardware.\nI was lucky and bought my two 3090s for about $750 each,<sup>[4](#fn:4)</sup> and I write them off over three years.\nThe GPUs only save money while they do work I would otherwise pay an API for, and while they sit idle they save nothing.\nSo the question is how many hours a day they have to be busy before they pay for themselves.[5](#fn:5)\n\nTo make this fair to open models, imagine an open model that is exactly as efficient as Luna: the same quality from the same number of tokens, sold at Luna’s price.\n\nWith free power, the GPUs pay for themselves if they are busy 4 hours a day, every day for three years. At 18 cents per kWh, they need 8.5. That uses what I paid for the cards; at today’s price of about $1,500 each, the 8.5 hours become 14.\n\nAnd API prices keep falling. GPT-6 Luna costs 58% less per output token than GPT-5.6 Luna did in July. If the price halves again, no amount of use pays off at 18 cents per kWh. If it halves twice, the API costs less than my electricity alone.\n\nEight and a half hours a day sounds like a lot, but I run agents for longer than that, just not on Qwen. When Claude Opus 4.6 came out in February, I was perfectly happy with it. I thought it was all I would ever need, and I could not have imagined how much better models would get in half a year. Qwen3.8 27B now scores about the same as Opus 4.6, and I would no longer accept it for coding. Until GPT-6 Astra came out at the start of September, my go-to was GPT-5.6 Sol, which scores ten points higher than Qwen. I then used Astra until Claude Opus 5.5 came out less than three weeks later. Now I don’t even accept what was considered the best model a month ago.\n\nThe models that fit on my 3090s keep improving, but they stay about half a year behind the frontier, and my standard moves with the frontier.\n\nFor me, this is a hobby, and a hobby is allowed to be inefficient. It gets worse once you try to use self-hosted models seriously, say for a team of ten developers.\n\nIf you stay fully self-hosted and want fast responses at peak, you have to buy hardware for the busiest hour of the year, not for the average one. The rest of the time, it sits idle.\n\nIn this made-up but realistic week, the team uses 17% of what it paid for, so every token costs about six times more than it would at full load. You also need a spare GPU for when one dies, and someone who gets paged when it does. You could queue work or send the peaks to an API, but then you are paying for an API anyway.\n\nAn API provider has the opposite situation. It serves thousands of customers across every time zone, so its load curve is much flatter than yours. It fills the nights with discounted batch jobs (Luna’s batch tier is half price) and training runs. And more concurrent requests mean bigger batches, which is where the efficiency from section 3 comes from.\n\nAnother argument I hear is that API prices are subsidized by investors and will go up once those investors want their money back. I don’t think that holds for inference.\n\nAgentPerf also reports how many tokens each GPU serves. On the GB300 rack, one GPU serving DeepSeek V4 Pro handles about 4 million output tokens and 530 million input tokens per hour. Most of that input is conversation history that coding agents send again with every step. Providers keep it in a cache and charge less than 1% of the normal input price for it.\n\nAt [DeepSeek V4 Pro’s list prices](https://artificialanalysis.ai/models/deepseek-v4-pro-0424), that GPU brings in at least $5.70 per hour, even if every input token is billed at the cache price.\nIf 5% of the input misses the cache, it brings in about $17.\nRenting a B200, the closest GPU with a [published rental price](https://artificialanalysis.ai/hardware-inference-stack/datacenter), costs $3.50 to $5.90 per hour from smaller cloud providers, and that price already includes their profit.\n\nSo the tokens pay for the hardware that serves them, even in the worst case. This leaves out costs like staff and training, so it does not tell you whether a lab makes money overall.\n\nThe money goes to training, research, free users, and flat-rate subscriptions for heavy users like me.\nWhen I wrote that I used [$10,000 worth of API tokens for $200](https://nijho.lt/post/agentic-coding/), that was list price, not what those tokens cost to serve.\n\nAnd prices go down, not up. This is what the same level of intelligence has cost this year:\n\nFrom Claude Opus 4.6 in February to GPT-6 Luna in September, the output price for this level of intelligence dropped by a factor of 50. Labs are in a race to the bottom on price.\n\nSome people say APIs are cheap because the provider trains on your data or sells it.\nThat is why I only used prices from providers that OpenRouter lists as [zero data retention](https://openrouter.ai/docs/features/zdr) (ZDR).\n\n| Model | Provider | ZDR | Output ($ per million tokens) | \n|---|---|---|---|\n| GPT-6 Luna | OpenAI | no | 0.50 | \n| GPT-6 Luna | Azure | yes | 0.50 | \n| Qwen3.8 27B | AkashML (FP8) | yes | 1.78 | \n| Qwen3.8 27B | DeepInfra (full precision) | yes | 1.88 | \n| Qwen3.8 27B | Alibaba | no | 2.55 | \n| Qwen3.8 27B | Cloudflare | no | 3.20 | \n\nAzure serves GPT-6 Luna with zero data retention at the same price as OpenAI.\nFor Qwen3.8 27B, the cheapest endpoints are all ZDR, and the endpoints without ZDR cost more.\nThe `:free` tier is where you pay with your data.\n\nIf you compare running Qwen3.8 27B yourself with paying for Qwen3.8 27B through an API, self-hosting wins.\n\nDeepInfra, the cheapest full-precision ZDR provider, charges $619 for one benchmark run, while my electricity costs $83. That compares the full model with my quantized copy, so the gap overstates my advantage by whatever quantization costs in quality. My machine pays for itself if it is busy 2.4 hours a day, or 1.5 hours with free power (the first group in the chart in section 4).\n\nProviders charge a lot for a dense 27B model, because every token runs through all 27 billion parameters. Sparse open models, which use only a small part of their weights for each token, can be much cheaper to serve. If you specifically need Qwen and use it heavily, self-hosting it pays off. But you don’t need Qwen to get answers of Qwen’s quality. Luna gives you those for $67, with zero data retention.\n\nFirst, it is fun.\nI like knowing how the whole stack works, from the drivers in [my NixOS configuration](https://nijho.lt/post/llama-nixos/) to how the layers are split between my two GPUs.\n\nSecond, nobody can take it away.\nArtificial Analysis already lists GPT-5.6 Luna as deprecated, less than three months after its release.\nMy Qwen weights will still be on my disk in ten years, and no company can change their license or [close their build system](https://nijho.lt/post/truenas-to-nixos/) on me.\n\nThird, some data should not leave my house.\nI would never send 200 GB of email, my messages, and my location history to an API, zero data retention or not.\nThat is what [my local AI projects](https://nijho.lt/post/local-ai-journey/) are for.\n\nThose are good reasons. Saving money is not one of them.\n\nFor each model, Artificial Analysis publishes how many input and output tokens the whole index took, and how many of the input tokens were read from a cache. I multiplied those by each provider’s prices, with cached input at the cache price. For Qwen that is 198 million output tokens and 820 million input tokens that miss the cache. For the bars at home, the same tokens divided by the machine’s speed give the run time, with the uncached input processed at the 1,600 tokens per second I measured on my 3090s and an estimated 2,000 on the RTX PRO 6000, and the run time times the machine’s power and the price per kWh gives the electricity. [↩︎](#fnref:1)\n\nThe RTX PRO 6000 has the same memory bandwidth as an RTX 5090, and at full precision every token reads three times as many bytes as in Q4_K_M. Scaling Artificial Analysis’s 5090 measurement gives about 48 tokens per second. I assume 600 W for the whole machine. [↩︎](#fnref:2)\n\nI measured the model as I normally run it: vLLM with an AutoRound INT4 quant and DFlash2 speculative decoding, split across both cards. Single 1,500-token answers came out at 105 to 109 tokens per second, and reading a 22,000-token prompt ran at 1,600 tokens per second. Sending 2, 4, or 8 requests at once did not raise the total above about 110 tokens per second. The two GPUs drew 540 W together, each at its 270 W power limit; the 700 W adds my estimate for the rest of the machine, since I could not measure the CPU. The setup follows the recipes from [club-3090](https://github.com/noonghunna/club-3090), a community project that tunes LLM serving for RTX 3090s and publishes measured numbers, so I think it is about as fast as these cards get. Its benchmarks for this configuration at a similar power limit match mine: about 110 tokens per second for prose, and up to about 195 for code, where speculative decoding guesses more tokens right. Most of Qwen’s output on the benchmark is reasoning, which runs at the prose speed. [↩︎](#fnref:3)\n\nThe 3090 is still the value king for VRAM per dollar, and $750 badly understates what mine are worth. That is what I paid more than a year ago; today a used 3090 sells for about $1,500. Even at that price it costs $62.50 per GB of VRAM, the same as an RTX 5090 at its $2,000 launch price, which is not what a 5090 sells for today. The right number for this calculation is what I could sell my cards for, and at $1,500 each the break-even points rise by about two thirds: at 18 cents per kWh, the same-model case moves from 2.4 to 4.0 hours per day, and the Luna-efficient case from 8.5 to 14.2 hours per day. [↩︎](#fnref:4)\n\nWhile the cards are busy, they save what the API would have charged for the same work, minus the electricity at 700 W. While they are idle, the machine still draws about 150 W, of which the two GPUs with the model loaded take 85 W. The break-even point is the number of busy hours per day at which those savings cover the price of the cards, written off over three years, plus the idle power. [↩︎](#fnref:5)", "url": "https://wpnews.pro/news/self-hosting-ai-does-not-save-money-and-i-do-it-anyway", "canonical_source": "https://nijho.lt/post/self-hosting-ai-is-not-cheaper/", "published_at": "2026-10-02 00:00:00+00:00", "updated_at": "2026-10-02 19:38:53.558292+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-infrastructure", "ai-tools"], "entities": ["Qwen3.8 27B", "GPT-6 Luna", "Claude Opus 4.6", "OpenAI", "Anthropic", "Artificial Analysis", "OpenRouter", "RTX 3090"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/self-hosting-ai-does-not-save-money-and-i-do-it-anyway", "markdown": "https://wpnews.pro/news/self-hosting-ai-does-not-save-money-and-i-do-it-anyway.md", "text": "https://wpnews.pro/news/self-hosting-ai-does-not-save-money-and-i-do-it-anyway.txt", "jsonld": "https://wpnews.pro/news/self-hosting-ai-does-not-save-money-and-i-do-it-anyway.jsonld"}}