{"slug": "custom-models-in-oh-my-pi-vllm-llama-cpp-sglang-and-more", "title": "Custom Models in Oh My Pi: vLLM, llama.cpp, SGLang and More", "summary": "Oh My Pi (omp) 18.2.7 changed how custom local models are configured, requiring users to rename the provider from `local` to an unused name and add `qwenTemplateReasoningEffort: true` to a model's `compat` block so Qwen 3.8 still receives an effort level instead of defaulting to `xhigh`. The release documents working `~/.omp/agent/models.yml` configs for vLLM, llama.cpp, LM Studio, Ollama, SGLang, Lemonade, ninfer and gateways such as Bifrost, with ports 8080, 1234, 11434, 30000 and 13305 and a 262144-token context window for the 27B Qwen entry.", "body_md": "# Custom Models in Oh My Pi: vLLM, llama.cpp, SGLang and More\n\nCustom models let you run Oh My Pi (omp) on local models, or on models from providers it doesn’t support out\nof the box. omp reads them from `~/.omp/agent/models.yml`. Below are working configs for the popular\ninference servers and for a gateway. After those come model roles, an allowlist that keeps omp off hosted\nmodels, thinking effort, fallbacks, and a way to log what omp sends.\n\nThere have been recent changes to omp. If you copied the config from my\n[tuning post](https://doug.sh/posts/tuning-a-local-coding-agent-oh-my-pi/), update it for\n[omp 18.2.7 and later](#what-changed-in-1827):\n\n- Rename the provider from `local` to a name omp doesn’t already use, and update the`modelRoles` entries to\nmatch. omp now has its own`local` provider for small on-device models, so your Qwen model under that name\nstops resolving.\n- Add `qwenTemplateReasoningEffort: true` to the model’s`compat` block. Without it, omp stops sending Qwen\n3.8 an effort level, and the model’s chat template picks`xhigh` every time.\n\n## Inference servers [🔗](#inference-servers)\n\n### vLLM [🔗](#vllm)\n\nFor a single vLLM server, this is all you need in `models.yml`:\n\n```\nproviders:\n  vllm:\n    baseUrl: http://192.168.1.20:8000/v1\n    auth: none\n    compat:\n      extraBody:\n        thinking_token_budget: 8192  # needs server support\n    modelOverrides:\n      qwen3.8-27b:                   # the name vLLM serves\n        maxTokens: 32768\n```\n\nomp gets the model list and context window from vLLM itself, and sends Qwen 3.8’s effort level without any extra setting.\n\n### llama.cpp, LM Studio and Ollama [🔗](#llamacpp-lm-studio-and-ollama)\n\nomp finds these on its own when they’re running locally on their default ports (8080, 1234, and 11434). For\none on another machine, set `LLAMA_CPP_BASE_URL`, `LM_STUDIO_BASE_URL`, or `OLLAMA_HOST`, or add the URL to\n`models.yml`:\n\n```\nproviders:\n  llama.cpp:\n    baseUrl: http://192.168.1.20:8080\n    api: openai-responses\n    auth: none\n    discovery:\n      type: llama.cpp\n```\n\nllama.cpp and LM Studio get Qwen 3.8’s effort level automatically, like vLLM.\n\n### SGLang, Lemonade, ninfer and the rest [🔗](#sglang-lemonade-ninfer-and-the-rest)\n\nThese have no built-in provider, but they speak the OpenAI API, so declare one and let omp read the model list:\n\n```\nproviders:\n  sglang:\n    baseUrl: http://192.168.1.20:30000/v1\n    api: openai-completions\n    auth: none\n    discovery:\n      type: openai-models-list\n    compat:\n      qwenTemplateReasoningEffort: true  # for Qwen 3.8\n```\n\nSGLang listens on port 30000, Lemonade on 13305, and ninfer on 8080, all under `/v1`. ninfer ignores\n`thinking_token_budget`, so set its budget with `--default-thinking-budget` when you start\nit. The [omp-ninfer](https://github.com/alphastorm/omp-ninfer) project has a tested omp setup for it.\n\n## Gateways and hand-listed models [🔗](#gateways-and-hand-listed-models)\n\nA gateway, or two servers of the same kind, needs a provider of your own too. Give it a name omp doesn’t\nalready use. I name mine after the gateway or machine behind it, and list the models by hand. Here’s the\n27B entry from my [Bifrost](https://github.com/maximhq/bifrost) gateway:\n\n```\nproviders:\n  bifrost:\n    baseUrl: http://192.168.1.10:8080/v1\n    api: openai-completions\n    apiKey: MY_GATEWAY_API_KEY               # env var (else literal)\n    headers:\n      x-bf-passthrough-extra-params: \"true\"  # or extraBody is dropped\n    models:\n      - id: rtx3090/qwen3.8-27b              # Bifrost's provider/model\n        name: qwen3.8-27b\n        contextWindow: 262144\n        maxTokens: 32768\n        reasoning: true\n        input: [text]\n        cost: {input: 0, output: 0, cacheRead: 0, cacheWrite: 0}\n        thinking:\n          mode: effort\n          efforts: [low, medium, xhigh]\n          defaultLevel: medium\n        compat:\n          qwenTemplateReasoningEffort: true  # needed since 18.2.7\n          extraBody:\n            thinking_token_budget: 8192\n```\n\nBifrost picks the backend from the part of the id before the slash. `rtx3090/` is a box with two RTX 3090s,\nand `strixhalo/` is the mini PC running Qwen3.8 Flash Next.\n\nIf `apiKey` starts with `!`, omp runs it as a command, which works with a password manager like 1Password:\n`\"!op read op://dev/gateway/key\"`.\n\n## Model roles [🔗](#model-roles)\n\n`modelRoles` in `~/.omp/agent/config.yml` decides which model does which job. `default` is the main agent\nand `task` runs subagents. A few more cover jobs like plan mode and commit messages. You can also set them from the Roles view in `/model`. A `:level` suffix sets the effort for that\nrole, and I run subagents at `low` and plan mode at `xhigh`:\n\n```\nmodelRoles:\n  default: bifrost/strixhalo/qwen3.8-flash-next:medium\n  plan: bifrost/strixhalo/qwen3.8-flash-next:xhigh\n  task: bifrost/rtx3090/qwen3.8-27b:low\n  smol: bifrost/rtx3090/qwen3.8-27b:low\n```\n\n## Keeping omp on your own models [🔗](#keeping-omp-on-your-own-models)\n\nIf a role’s model doesn’t resolve, or `models.yml` doesn’t parse, omp doesn’t stop. It falls back to the\ndefault model of a known provider it can use, and failing that, the first model it can use at all. It can\nuse any provider that needs no key or that it has a key for. Keys come from logins, from a few dozen\nenvironment variables, some of which you probably set for other tools, like `HF_TOKEN` or\n`AZURE_OPENAI_API_KEY`, and from `.env` files, including one in the project you’re working in. With an\nAWS Bedrock token in the environment, the tuning post’s `local` config sent my test prompt to\nClaude Opus 5.5 on Bedrock.\n\nAn allowlist in `config.yml` prevents that:\n\n```\nenabledModels:\n  - \"bifrost/*\"\n```\n\nNow omp only starts on a `bifrost` model, and stops at startup if none of them resolves. A project’s `.omp/config.yml` replaces this list\nrather than adding to it, so a project with its own list needs `bifrost/*` in it too.\n\n## Thinking effort [🔗](#thinking-effort)\n\n`efforts` lists the levels the model accepts, and `defaultLevel` is the one omp uses when a role has no\nsuffix. Qwen 3.8 takes `low`, `medium`, and `xhigh`.\n\nKeep the `thinking_token_budget` too. Together with the effort level, it’s what fixed the\n[five-minute turns in the tuning post](https://doug.sh/posts/tuning-a-local-coding-agent-oh-my-pi/#the-agent-thought-for-five-minutes-before-doing-anything),\nand I’ve since raised it to 8,192 tokens without seeing the cap make the model worse at real work.\n\n## What the catalog fills in [🔗](#what-the-catalog-fills-in)\n\nIf you list a model by hand and leave a field out, omp copies it from its built-in catalog. A bare\n`qwen3.8-27b` ends up with a hosted price, image input, and a 65,536-token reply limit. The price only\nchanges omp’s cost estimate.\n\nomp also ignores a misspelled key and uses the catalog’s value, so `maxToken: 32768` gets you the catalog’s\nreply limit instead of yours. For your own providers, write out every field in the Bifrost example, and run\n`omp models <provider>` to verify.\n\n## Fallbacks [🔗](#fallbacks)\n\n`retry.fallbackChains` in `config.yml` says what to try when a model keeps failing. A key can be a role, a\nmodel, or `provider/*`:\n\n```\nretry:\n  fallbackChains:\n    bifrost/strixhalo/qwen3.8-flash-next:\n      - bifrost/rtx3090/qwen3.8-27b:medium\n    bifrost/rtx3090/qwen3.8-27b:\n      - bifrost/strixhalo/qwen3.8-flash-next:low\n```\n\nIf a model’s server is down, omp moves to the next one in its chain and, by default, switches back on its own\nlater. A hosted model can go in a chain if it’s also in `enabledModels`, but then omp can start on it when a role doesn’t resolve.\n\n## What changed in 18.2.7 [🔗](#what-changed-in-1827)\n\nBoth changes are in the [18.2.7 release notes](https://github.com/can1357/oh-my-pi/releases/tag/v18.2.7).\nThe `local` provider is under Added:\n\nAdded model-kind and grounded-search capability metadata, along with catalogs for local inference and search-engine models.\n\nThe effort change is under Changed:\n\nImproved model routing and thinking-policy handling for llama.cpp Qwen models, Bonsai lineage aliases, and custom provider names.\n\n## Logging what omp sends [🔗](#logging-what-omp-sends)\n\nWhen omp does something odd with a model, I print what it’s sending. This simple script just prints each request body, minus the messages and tools, and answers “hi”:\n\n``` python\nimport json\nfrom http.server import BaseHTTPRequestHandler, HTTPServer\n\ndef sse_chunk(delta, finish):\n    choice = {\"index\": 0, \"delta\": delta, \"finish_reason\": finish}\n    chunk = {\"object\": \"chat.completion.chunk\", \"choices\": [choice]}\n    return f\"data: {json.dumps(chunk)}\\n\\n\".encode()\n\nclass Handler(BaseHTTPRequestHandler):\n    def do_POST(self):\n        length = int(self.headers[\"Content-Length\"])\n        body = json.loads(self.rfile.read(length))\n        body.pop(\"messages\", None)\n        body.pop(\"tools\", None)\n        print(json.dumps(body, indent=2), flush=True)\n        self.send_response(200)\n        self.send_header(\"Content-Type\", \"text/event-stream\")\n        self.end_headers()\n        self.wfile.write(sse_chunk({\"content\": \"hi\"}, None))\n        self.wfile.write(sse_chunk({}, \"stop\"))\n        self.wfile.write(b\"data: [DONE]\\n\\n\")\n\nHTTPServer((\"127.0.0.1\", 18080), Handler).serve_forever()\n```\n\nCopy your model into a scratch `models.yml` under a provider named `test`, with\n`baseUrl: http://127.0.0.1:18080/v1`, `api: openai-completions` and `auth: none`. Add `enabledModels: [\"test/*\"]` to a scratch `config.yml` so nothing\ncan fall back to a hosted model. Put both files in one directory and point omp at it, using your model’s id:\n\n```\nmkdir -p /tmp/omp-test  # models.yml and config.yml go here\nPI_CODING_AGENT_DIR=/tmp/omp-test \\\n  omp -p --no-session --model test/qwen3.8-27b:medium \"Say hi.\"\n```\n\n*I had help with this one. Anthropic’s Claude helped me test omp’s releases against the logging server\nand draft this post. I read all the words, checked the results, and rewrote anything that sounded like a\nchatbot, so the mistakes are mine.*\n\n## Sources [🔗](#sources)\n\n- omp releases: [18.2.7](https://github.com/can1357/oh-my-pi/releases/tag/v18.2.7) and[18.3.0](https://github.com/can1357/oh-my-pi/releases/tag/v18.3.0)\n- omp docs at 18.3.0: [models.md](https://github.com/can1357/oh-my-pi/blob/v18.3.0/docs/models.md) ,[settings.md](https://github.com/can1357/oh-my-pi/blob/v18.3.0/docs/settings.md) , and[environment-variables.md](https://github.com/can1357/oh-my-pi/blob/v18.3.0/docs/environment-variables.md)\n- omp’s Qwen 3.8 effort rules: [`vllm` and `lm-studio`](https://github.com/can1357/oh-my-pi/blob/v18.3.0/packages/catalog/src/compat/rules/classes/qwen.kdl) and[`llama.cpp`](https://github.com/can1357/oh-my-pi/blob/v18.3.0/packages/catalog/src/compat/rules/providers/llama.cpp.kdl)\n- [omp’s built-in `local` provider](https://github.com/can1357/oh-my-pi/blob/v18.3.0/packages/catalog/src/compat/rules/providers/local.kdl)\n- [Bifrost (`provider/model` routing)](https://github.com/maximhq/bifrost)\n- Engine docs: [SGLang server arguments (port 30000)](https://docs.sglang.io/docs/advanced_features/server_arguments) ,[Lemonade’s OpenAI-compatible API (port 13305)](https://lemonade-server.ai/docs/api/openai/) , and[ninfer serving (port 8080, `--default-thinking-budget`)](https://github.com/Neroued/ninfer/blob/master/docs/serving.md)", "url": "https://wpnews.pro/news/custom-models-in-oh-my-pi-vllm-llama-cpp-sglang-and-more", "canonical_source": "https://doug.sh/posts/oh-my-pi-custom-models/", "published_at": "2026-09-23 00:00:00+00:00", "updated_at": "2026-09-25 02:00:43.227634+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "developer-tools", "ai-infrastructure"], "entities": ["Oh My Pi", "vLLM", "llama.cpp", "SGLang", "LM Studio", "Ollama", "Bifrost", "Qwen 3.8"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/custom-models-in-oh-my-pi-vllm-llama-cpp-sglang-and-more", "markdown": "https://wpnews.pro/news/custom-models-in-oh-my-pi-vllm-llama-cpp-sglang-and-more.md", "text": "https://wpnews.pro/news/custom-models-in-oh-my-pi-vllm-llama-cpp-sglang-and-more.txt", "jsonld": "https://wpnews.pro/news/custom-models-in-oh-my-pi-vllm-llama-cpp-sglang-and-more.jsonld"}}