{"slug": "this-is-my-opencode-setup-for-local-models-mainly-using-a-customized-llama-cpp", "title": "This is my OpenCode setup for local models, mainly using a customized llama.cpp configuration.", "summary": "A developer detailed their OpenCode setup for running local AI models, centered on a customized llama.cpp configuration. The setup leverages reasoning-effort levels for models like Qwen 3.8, DeepSeek V4, and Glimmer, and applies token budgets for models without explicit levels. The configuration includes a llama-server command with flags for reasoning budgets and context preservation, plus an OpenCode configuration file for integrating local llama.cpp and vLLM providers.", "body_md": "This is my OpenCode setup for local models, mainly using a customized llama.cpp configuration.\n\nFor Qwen 3.8, DeepSeek V4, and Glimmer, the models are already trained to support reasoning effort levels. Depending on the model, these may be exposed as `low`\n\n, `medium`\n\n, `high`\n\n, `xhigh`\n\n, or as `low`\n\n, `high`\n\n, and `max`\n\n.\n\nFor models that support reasoning but were not trained with explicit reasoning-effort levels, such as the Qwen 3.5 and 3.6 variants, I use a token budget to limit the amount of reasoning.\n\nAlthough Qwen 3.8 has built-in reasoning-effort levels, I still apply a maximum reasoning-token cap for each effort level.\n\nAlthough the llama.cpp CLI flags specify `preserve_thinking`\n\nand a default reasoning budget, these can still be overridden through the API, so this works fine for my setup.\n\nYes, there is also an `xhigh-no-preserve`\n\nvariant. In this mode, the model uses its reasoning as a scratchpad without preserving it in the conversation history. I use this when I do not want the reasoning output to unnecessarily consume the context window.\n\nMost of the time, I use `low`\n\nreasoning.\n\n```\n/home/USER/llama.cpp/llama-server \\\n  --model /mnt/d/MODEL_STORE/LLM_SETUP/Qwen3.8-27B/Qwen3.8-27B-UD-Q5_K_XL.gguf \\\n  --alias Qwen3.8-27B \\\n  --host 0.0.0.0 \\\n  --chat-template-file /mnt/d/MODEL_STORE/LLM_SETUP/Qwen3.8-27B/chat_template.jinja \\\n  --no-context-shift \\\n  --metrics \\\n  --kv-unified \\\n  --cache-ram 16384 \\\n  --ctx-size 81920 \\\n  --port 8001 \\\n  --cache-type-k q8_0 \\\n  --cache-type-v q8_0 \\\n  --flash-attn on \\\n  --temp 1.0 \\\n  --top-p 0.95 \\\n  --top-k 20 \\\n  --min-p 0.0 \\\n  --presence-penalty 0.0 \\\n  --repeat-penalty 1.0 \\\n  --jinja \\\n  --chat-template-kwargs '{\"preserve_thinking\": true}' \\\n  --spec-type draft-mtp \\\n  --spec-draft-n-max 2 \\\n  --reasoning-budget 8192 \\\n  -np 1 \\\n  -ub 256\n```\n\n[https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)\n\n`nvim ~/.config/opencode/opencode.json`\n\n```\n{\n  \"$schema\": \"https://opencode.ai/config.json\",\n  \"plugin\": [\n    \"@tarquinen/opencode-dcp@latest\"\n  ],\n  \"provider\": {\n    \"openai\": {\n      \"options\": {\n        \"headerTimeout\": 60000,\n        \"timeout\": 600000,\n        \"chunkTimeout\": 60000\n      }\n    },\n    \"local-vllm\": {\n      \"npm\": \"@ai-sdk/openai-compatible\",\n      \"name\": \"vLLM (Local)\",\n      \"options\": {\n        \"baseURL\": \"http://localhost:8000/v1\"\n      },\n      \"models\": {\n        \"cyankiwi/Nemotron-Orchestrator-8B-AWQ-4bit\": {\n          \"name\": \"cyankiwi/Nemotron-Orchestrator-8B-AWQ-4bit\",\n          \"max_tokens\": 40960\n        }\n      }\n    },\n    \"local-llamacpp\": {\n      \"npm\": \"@ai-sdk/openai-compatible\",\n      \"name\": \"llama.cpp (Local)\",\n      \"options\": {\n        \"baseURL\": \"http://localhost:8001/v1\"\n      },\n      \"models\": {\n        \"Qwen3.6-35B\": {\n          \"name\": \"Qwen3.6-35B\",\n          \"max_tokens\": 131072,\n          \"modalities\": {\n            \"input\": [\n              \"image\",\n              \"text\"\n            ],\n            \"output\": [\n              \"text\"\n            ]\n          },\n          \"variants\": {\n            \"none\": {\n              \"chat_template_kwargs\": {\n                \"enable_thinking\": false\n              }\n            },\n            \"low\": {\n              \"reasoning_budget_tokens\": 512\n            },\n            \"medium\": {\n              \"reasoning_budget_tokens\": 2048\n            },\n            \"xhigh\": {\n              \"reasoning_budget_tokens\": 8192\n            },\n            \"xhigh-no-preserve\": {\n              \"reasoning_budget_tokens\": 8192,\n              \"chat_template_kwargs\": {\n                \"preserve_thinking\": false\n              }\n            }\n          }\n        },\n        \"Qwen3.8-27B\": {\n          \"name\": \"Qwen3.8-27B\",\n          \"max_tokens\": 81920,\n          \"modalities\": {\n            \"input\": [\n              \"image\",\n              \"text\"\n            ],\n            \"output\": [\n              \"text\"\n            ]\n          },\n          \"variants\": {\n            \"none\": {\n              \"reasoningEffort\": \"none\"\n            },\n            \"low\": {\n              \"reasoningEffort\": \"low\",\n              \"reasoning_budget_tokens\": 512\n            },\n            \"medium\": {\n              \"reasoningEffort\": \"medium\",\n              \"reasoning_budget_tokens\": 2048\n            },\n            \"xhigh\": {\n              \"reasoningEffort\": \"xhigh\",\n              \"reasoning_budget_tokens\": 8192\n            },\n            \"xhigh-no-preserve\": {\n              \"reasoningEffort\": \"xhigh\",\n              \"reasoning_budget_tokens\": 8192,\n              \"chat_template_kwargs\": {\n                \"preserve_thinking\": false\n              }\n            }\n          }\n        },\n        \"Muse-Glimmer-30B\": {\n          \"name\": \"Muse-Glimmer-30B\",\n          \"max_tokens\": 81920,\n          \"modalities\": {\n            \"input\": [\n              \"image\",\n              \"text\"\n            ],\n            \"output\": [\n              \"text\"\n            ]\n          },\n          \"variants\": {\n            \"low\": {\n              \"reasoningEffort\": \"low\",\n              \"reasoning_budget_tokens\": 512\n            },\n            \"medium\": {\n              \"reasoningEffort\": \"medium\",\n              \"reasoning_budget_tokens\": 2048\n            },\n            \"high\": {\n              \"reasoningEffort\": \"high\",\n              \"reasoning_budget_tokens\": 4096\n            },\n            \"xhigh\": {\n              \"reasoningEffort\": \"xhigh\",\n              \"reasoning_budget_tokens\": 8192\n            }\n          }\n        },\n        \"DeepSeek-V4-Flash-0731\": {\n          \"name\": \"DeepSeek-V4-Flash-0731\",\n          \"max_tokens\": 81920,\n          \"variants\": {\n            \"low\": {\n              \"reasoningEffort\": \"low\",\n              \"reasoning_budget_tokens\": 512\n            },\n            \"high\": {\n              \"reasoningEffort\": \"high\",\n              \"reasoning_budget_tokens\": 4096\n            },\n            \"max\": {\n              \"reasoningEffort\": \"max\",\n              \"reasoning_budget_tokens\": 8192\n            }\n          }\n        },\n        \"Omnicoder-2-9B\": {\n          \"name\": \"Omnicoder-2-9B\",\n          \"max_tokens\": 131072,\n          \"modalities\": {\n            \"input\": [\n              \"image\",\n              \"text\"\n            ],\n            \"output\": [\n              \"text\"\n            ]\n          },\n          \"variants\": {\n            \"none\": {\n              \"chat_template_kwargs\": {\n                \"enable_thinking\": false\n              }\n            },\n            \"low\": {\n              \"reasoning_budget_tokens\": 512\n            },\n            \"medium\": {\n              \"reasoning_budget_tokens\": 2048\n            },\n            \"xhigh\": {\n              \"reasoning_budget_tokens\": 8192\n            },\n            \"xhigh-no-preserve\": {\n              \"reasoning_budget_tokens\": 8192,\n              \"chat_template_kwargs\": {\n                \"preserve_thinking\": false\n              }\n            }\n          }\n        },\n        \"Ornith-9B\": {\n          \"name\": \"Ornith-9B\",\n          \"max_tokens\": 131072,\n          \"modalities\": {\n            \"input\": [\n              \"image\",\n              \"text\"\n            ],\n            \"output\": [\n              \"text\"\n            ]\n          },\n          \"variants\": {\n            \"none\": {\n              \"chat_template_kwargs\": {\n                \"enable_thinking\": false\n              }\n            },\n            \"low\": {\n              \"reasoning_budget_tokens\": 512\n            },\n            \"medium\": {\n              \"reasoning_budget_tokens\": 2048\n            },\n            \"xhigh\": {\n              \"reasoning_budget_tokens\": 8192\n            },\n            \"xhigh-no-preserve\": {\n              \"reasoning_budget_tokens\": 8192,\n              \"chat_template_kwargs\": {\n                \"preserve_thinking\": false\n              }\n            }\n          }\n        }\n      }\n    },\n    \"local-ninfer\": {\n      \"npm\": \"@ai-sdk/openai-compatible\",\n      \"name\": \"Ninfer (Windows)\",\n      \"options\": {\n        \"baseURL\": \"http://localhost:10909/v1\"\n      },\n      \"models\": {\n        \"qwen3.8-27b\": {\n          \"name\": \"Qwen3.8-27B\",\n          \"max_tokens\": 81920,\n          \"variants\": {\n            \"none\": {\n              \"reasoningEffort\": \"none\"\n            },\n            \"low\": {\n              \"reasoningEffort\": \"low\"\n            },\n            \"medium\": {\n              \"reasoningEffort\": \"medium\"\n            },\n            \"xhigh\": {\n              \"reasoningEffort\": \"xhigh\"\n            }\n          }\n        }\n      }\n    }\n  }\n}\n```\n\n- 3090\n- WSL2 Ubuntu 24.04\n- 6.18.33.2-microsoft-standard-WSL2\n- E5 2690v4\n- 96G RAM (For WSL2), Total 128G, DDR4 2400 ECC\n\n- Qwen 3.8 27B KV Q8:Q8 81920 : 30 tok/s\n- Qwen 3.6 35B A3B KV Q8:Q8 180224: 75 tok/s\n- DSv4 Flash KV Q8:Q8 65536: 6 tok/s", "url": "https://wpnews.pro/news/this-is-my-opencode-setup-for-local-models-mainly-using-a-customized-llama-cpp", "canonical_source": "https://gist.github.com/komikndr/b17955e1a80ce6ede9a3115f16216bc5", "published_at": "2026-08-18 04:12:42+00:00", "updated_at": "2026-08-18 04:41:18.332183+00:00", "lang": "en", "topics": ["developer-tools", "large-language-models", "ai-infrastructure", "ai-tools"], "entities": ["OpenCode", "llama.cpp", "Qwen 3.8", "DeepSeek V4", "Glimmer", "vLLM", "Hugging Face", "froggeric"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/this-is-my-opencode-setup-for-local-models-mainly-using-a-customized-llama-cpp", "markdown": "https://wpnews.pro/news/this-is-my-opencode-setup-for-local-models-mainly-using-a-customized-llama-cpp.md", "text": "https://wpnews.pro/news/this-is-my-opencode-setup-for-local-models-mainly-using-a-customized-llama-cpp.txt", "jsonld": "https://wpnews.pro/news/this-is-my-opencode-setup-for-local-models-mainly-using-a-customized-llama-cpp.jsonld"}}