I am optimistic about the potential of running large language models on local systems. Local LLMs have many benefits, and I run them frequently, mostly for chat or information lookups.
But for more complex or difficult tasks, like generating complicated Rust functions, I feel that cloud-based LLMs are a better deal. In trying to generate code on my local PC, I was successful but it took up to 30 minutes or an hour of 100% CPU and 100% GPU work.
I estimated my cost in electricity and it came out to about $0.02 to $0.04 for the generations. This might make sense in some cases, but running the same generations on OpenRouter with a modern Flash model (Gemini 3.8 Flash, GLM 5.3 Flash, Claude Sonnet 5.5) is about the same price, and usually finishes within 2 minutes.
So cloud-based AI is both faster, cheaper, and also higher quality (based on a larger model). It does mean that your data is no longer "private" and if you are doing something questionable a model might refuse to generate text. There are drawbacks with both local and cloud, but often cloud is a better deal for code generation.