Problems with Local LLMs
A developer's comparison found that generating complicated Rust functions on a local LLM took 30 to 60 minutes of 100% CPU and 100% GPU usage at an estimated electricity cost of $0.02 to $0.04, while …
A developer's comparison found that generating complicated Rust functions on a local LLM took 30 to 60 minutes of 100% CPU and 100% GPU usage at an estimated electricity cost of $0.02 to $0.04, while …
A developer reported that Claude Sonnet 5.5 generated two Rust functions for about $0.06 after DeepSeek V4.1 Flash produced hard-to-read code and a slight performance regression. The developer said So…
A developer replaced the Rust `clap` crate with LLM-generated argument-parsing code written by Gemini 3.8 Flash via OpenRouter, cutting the compiled executable size by 300 kB and argument parsing time…
A developer reported that Claude Opus 5.5 generated correct Rust code from a spec for $0.09 per final generation, improving the code by switching lookup table values from u32 to u8 for slightly faster…
Spec-driven development on OpenRouter cut the cost of evaluating a Markdown prompt of up to 180 lines and generating a function to between 1 and 7 cents, according to the author's experiments, far bel…
Xiaomi's MiMo-V2.6 Flash, a scaled-down version of the Pro model, rewrote a Rust codebase to replace lookup tables with direct match statements in 17 seconds at a cost of $0.0005 via OpenRouter, produ…
A developer used OpenRouter to run GLM 5.3 Flash on two Rust functions generated from written specs, paying about $0.03 total for the work. The developer reported the model coded the functions well, w…
GPT-6-Sol generated correct Rust code from two Markdown specifications in under a minute for about $0.07, according to a developer's first-person account of spec-driven development. The same two code …
A developer reports that writing detailed Markdown specification files and having an LLM generate code from them produces reliable output, using the Qwen 3.8 27B model (Bonsai Ternary 2) for functions…
A developer reported that a local model, Ternary Bonsai 2, generated the parse_prompt function returning ParsedPromptInfo from a 180-line Markdown specification in about 30 minutes, producing 300-400 …
A developer who implemented a simple MCP server for file creation, deletion, and viewing reported that the LLM added little value, since calling the program executable directly was easier. Using Bonsa…
A user improved the agentic coding performance of Ternary Bonsai 2, a Qwen 3.8 27B-based model, when run locally through llama-cpp by switching its thinking setting from "xhigh" to medium via the chat…
A tester ran Bonsai 2 27B, a ternary model that stores each weight as one of three values (1, 0, -1), locally on a 12GB Nvidia GPU using a special build of llama-cpp from PrismML's GitHub. The 7.2 GB …
A developer's side-by-side test found that Google's Gemini 3.8 Flash cloud model produced a faster Rust parsing function than the local Gemma 4 model (gemma-4-26B_q4_0-it.gguf), cutting starts_with ca…
A user tested Qwen 3.5 0.8B (Qwen3.5-0.8B-Q4_K_M.gguf) locally and found it works on their MCP server, with a temperature of 0 giving the most consistent results for agentic function calling despite d…
A developer expressed doubts about agentic AI coding tools such as OpenCode and Pi, saying they have had "mixed results at best" and get "better results" from chatbot-style development. The developer …
A developer reports that prompting a chat LLM for individual functions and pasting the output into code can work better than agentic harnesses for some situations, and that a local model such as Gemma…
A developer got the 230M-parameter LiquidAI LFM2.5-230M-Q5_K_M.gguf model working with a local MCP server after switching from 4-bit to 5-bit quantization, following an earlier setup that ran the 350M…
A developer comparing open-source agentic coding harnesses Pi and OpenCode found Pi faster and more responsive, with better interface updates and no AI subscription sales, giving Pi the edge for now. …
Laguna XS 2.1, a small local AI model, successfully refactored a Rust function on a 12 GB Nvidia GPU, according to a developer's account. The user reported minimal errors and task completion, but note…