Show HN: Run Claude Frontend with Non-Anthropic Models A developer released a local proxy that translates Anthropic's Messages API to OpenAI-compatible backends, letting the stock Claude Code CLI run against OpenAI keys or on-device models such as Ollama's qwen2.5:7b-instruct. The proxy, run via `node supervise.mjs` on 127.0.0.1:8123, was verified end-to-end with the unmodified CLI on a full agent turn using 221 tools, with logs showing model=claude-opus-4-8->gpt-5.6-sol. The accompanying desktop build repoints Anthropic's unpacked Claude Desktop app under a stock Electron runtime but requires users to supply their own licensed copy of the ~40 MB proprietary bundle, which is not redistributable. Claude Code and Claude Desktop are good agent harnesses. This repo lets you drive their agent layer with a backend of your choice — an OpenAI key, an on-device model on your own GPU , or stock Anthropic — by translating between the Anthropic and OpenAI APIs locally. The chat window is always the real claude.ai web app; only the agent sub-layer Claude Code is repointed. The desktop build's Code tab. The model picker lists Composite the cross-provider fallback chain plus every keyed provider's models as provider:model — Gemini, Mistral, nvidia, ollama-cloud, llm7, freetoken, cloudflare — each routing that turn to its own upstream through the proxy; the local usage dashboard summarises agent activity. Two independent things live here, and you may only want the first : | | What it is | What you need | |---|---|---| | 1. The proxy — openai-proxy/ https://github.com/mkornreich/llm-desktop-electron/blob/main/openai-proxy | A local server that speaks the Anthropic Messages API on the front and calls OpenAI or any OpenAI-compatible server on the back. Point any Anthropic-API client at it. | Node 22+ native fetch , and an OpenAI key or a local server like Ollama | | 2. The desktop build | Anthropic's Claude Desktop app, unpacked and run under a stock Electron runtime, with the proxy wired in, a provider switch, and a settings GUI | Linux or macOS, plus your own copy of Claude Desktop | The proxy is the generally useful part: it works with the ordinary claude CLI on any machine, and needs nothing from this repo's app/ directory. Everything below the Desktop build 2-the-desktop-build heading is specific to running Anthropic's Electron app. Licensing, read this first. app/ is Anthropic's unpacked proprietary bundle ~40 MB of their JavaScript, native binaries and resources . It is not redistributable, which is why this repository is private. For part 2 you must supply that directory from your own licensed Claude Desktop install — see Supplying the bundle supplying-the-bundle . Nothing in part 1 touches it. export OPENAI API KEY=sk-... or point at a local server; see below cd openai-proxy && node supervise.mjs listens on 127.0.0.1:8123, restarts on crash node proxy.mjs still works and is right for a one-off. Prefer the supervisor for anything you leave running: the proxy has taken itself down on a dropped upstream socket with nothing noticing for hours — the app, the launcher and every agent kept running against a closed port. The supervisor restarts it, logs why, and gives up loudly rather than looping if it cannot start. Point the stock Claude Code CLI at it: ANTHROPIC BASE URL=http://127.0.0.1:8123 ANTHROPIC API KEY=unused claude -p "hello" ANTHROPIC API KEY must be set to something because the CLI insists on it, but the proxy ignores it and authenticates to OpenAI with OPENAI API KEY . Verified end-to-end with the unmodified CLI: a full agent turn, 221 tools, tool calls executed and results fed back, with the proxy log showing model=claude-opus-4-8- gpt-5.6-sol . The CLI still says "Opus" because it hardcodes that in its own prompt; the proxy log and the ANTHROPIC BASE URL in the process environment are the ground truth. Set OPENAI BASE URL at any OpenAI-compatible server Ollama https://ollama.com , llama.cpp, LM Studio, vLLM and the proxy translates to it exactly as it does to OpenAI. A loopback base URL is treated as keyless — no OPENAI API KEY needed a placeholder bearer is sent, which local servers ignore : OPENAI BASE URL=http://127.0.0.1:11434/v1 OPENAI MODEL=qwen2.5:7b-instruct OPENAI API=chat \ node openai-proxy/proxy.mjs The desktop build wires this up as the local provider on-device-model-local , which also starts and tunes Ollama for you. Everything has a default; set only what you want to change. | Variable | Default | What it does | |---|---|---| | OPENAI API KEY | — | Required for a remote endpoint; optional for a loopback one . Also read from .openai-key see below . | | OPENAI MODEL | gpt-5.6-sol | Any model id. Names containing codex route to Responses automatically; anything else lands on Chat Completions unless OPENAI API=responses is set. | | OPENAI BASE URL | https://api.openai.com/v1 | Any OpenAI-compatible endpoint, including a local one. | | PORT | 8123 | | | OPENAI API | responses | responses or chat , overriding the name heuristic. Required for a non-codex remote model : this app sends 200+ tools and Chat Completions caps at 128. | | OPENAI REASONING EFFORT | max | The proxy steps down to the highest value your model accepts. | | OPENAI CLASSIFIER SAFETY MODEL | gpt-5.4-2026-03-05 | Model for Claude Code's auto-mode safety verdict. It has a 60-second deadline and denies the action when it expires, so this wants a fast model — see openai-proxy/README.md https://github.com/mkornreich/llm-desktop-electron/blob/main/openai-proxy/README.md for the measurements. | The full list — around thirty knobs, each documented with the bug that motivated it — is in openai-proxy/README.md https://github.com/mkornreich/llm-desktop-electron/blob/main/openai-proxy/README.md . Instead of environment variables, put KEY=VALUE lines in .openai-model at the repo root checked in, for non-secrets . The secret API key goes in its own gitignored file, .openai-key apiKey=… ; copy .openai-key.example to start . Precedence is env var → .openai-model → .openai-key → default. That precedence lives in exactly one place, openai-proxy/config.mjs https://github.com/mkornreich/llm-desktop-electron/blob/main/openai-proxy/config.mjs , which also produces the config hash below. To see what a launch would actually use: node openai-proxy/config.mjs the effective config, with the source of each value --hash prints just the config hash; --validate checks it and exits non-zero on an error. The key never appears in any of them — only a sha256: fingerprint of it. /health reports an identity, not just liveness: { "ok": true, "instance": "8f3c…", "pid": 82110, "configHash": "38d9c8b2…", "codeVersion": "a8e08f01…", "inflight": 0, "model": "gpt-5.6-sol", "…": "…" } - instance is a nonce generated at startup and also written to proxy-runtime.json . Only the process that generated it can serve it, so matching both proves ownership — which is what lets the launcher restart its proxy and refuse to touch anything else. - configHash covers every behaviour-affecting setting plus a fingerprint of the key. The launcher compares it and restarts a proxy serving stale settings, instead of the old behaviour where any answer on /health counted as healthy and a changed model silently did not apply. - codeVersion is a hash of the proxy sources, catching a process running older translation logic with current settings. - inflight is how many turns are being answered right now, so a restart can say what it would interrupt. Supervisor behaviour is tuned by environment variable; the defaults are meant to be left alone. PROXY MAX FAST FAILURES 8 bounds immediate restart attempts, PROXY RESTART BASE MS 500 and PROXY RESTART MAX MS 30000 set the backoff, and PROXY HEALTH EVERY MS 15000 / PROXY HEALTH TIMEOUT MS 5000 / PROXY HEALTH MISSES 3 control the watchdog that replaces a proxy holding the port without answering. An externally-killed proxy always restarts and never counts toward the give-up bound. | | Chat Completions | Responses | |---|---|---| | Tool limit | 128 129 → HTTP 400 | none observed — probed to 512 | | Reasoning controls | none | reasoning.effort , reasoning summaries | | Local-server support | universal | recent Ollama; not all local servers | Claude Code offers 27 tools; Claude Desktop offers over 200. On the chat surface the 128-tool cap is real, so the proxy keeps the essential tools read/write/edit/run/search/plan/web, plus renderers and fills the remainder in the agent's own order, logging what it dropped — a silently truncated tool list is indistinguishable from a model that just declined to use a tool. Text and tool calls in both directions, streaming and non-streaming, plus fixes for behaviour differences between the two model families: - Rendering — GPT models emit \ …\ / \ …\ for maths, which Claude's renderer shows literally. Rewritten to $…$ / $$…$$ , fence-aware so code samples stay verbatim. - Persistence — GPT models routinely end a turn to check in "If you want, I'll run that now" , which in an agent loop reads as a stall. An explicit persistence directive plus auto-continue when a turn announces an action and calls no tool. - Context overflow — Claude Code sizes its own auto-compaction from the model it thinks it is talking to, so it never fires. The proxy compacts and retries, including when the overflow arrives as a mid-stream event rather than an HTTP 400. - Empty and truncated turns — retried or resumed rather than surfaced as a blank reply. - Unsupported parameters — the CLI sends stop sequences , which some models reject. The proxy drops whichever parameter the API names and retries, so the next unsupported knob self-heals too. - Thinking — OpenAI reasoning summaries are mapped to Anthropic thinking blocks. Raw chain-of-thought is not available from the API at any setting. Images work. Anthropic {type:"image", source:{…}} blocks — base64 or url — become image url on Chat Completions and input image on Responses. If the model has no vision the picture is dropped and the question kept, with an honest note in its place, rather than failing the turn. Known gaps: /v1/messages/count tokens is estimated, and the cost figures the CLI prints are computed from Anthropic's price list and are meaningless when proxied. npm run eval run the corpus, diff against the frozen baseline npm run eval:freeze overwrite the baseline with what the code does NOW eval/ holds a synthetic, network-free corpus. Each case runs through the real proxy against a fake upstream, and what gets recorded is the payload the proxy decided to send: resolved model, tool count, whether a late tool survived, tool choice , hint injection, reasoning, verbosity, output cap, cache key, pricing tier. Every one of those decisions is settled before a token is generated, which is what makes it a baseline rather than a sample. npm test checks the frozen baseline on every run, so a behaviour change cannot arrive unnoticed inside an unrelated commit. Alongside it are invariants the baseline is not allowed to encode away — a classifier is never given tools, a safety verdict never resolves to the main model, an agent turn that merely quotes the classifier contract keeps its full catalogue. npm test the whole suite: proxy + settings + launcher scripts + eval baseline Most were written against a specific misbehaviour; the comments say which. Runs Anthropic's Claude Desktop @ant/desktop v1.24012.9, an Electron Forge + Vite build from its unpacked app.asar under a stock Electron runtime, with the proxy wired into the agent layer, a provider switch, and a settings GUI over the launcher's config. Runs on Linux and macOS . app/ is committed in this repository — Anthropic's code — and that is exactly why it is private and must not be published. node modules/ and app/node modules/ are not committed; you populate them from your own licensed install. To rebuild app/ from your own install: 1. Find app.asar in your Claude Desktop installation: - Linux: /usr/lib/claude-desktop/resources/ - macOS: /Applications/Claude.app/Contents/Resources/ 2. npx @electron/asar extract app.asar app/ 3. Merge the sibling app.asar.unpacked/node modules/ back in over app/node modules/ — the native addons @ant/claude-native , node-pty , and on macOS @ant/claude-swift live there as platform binaries and the app expects them beside their JS. On Linux, @ant/claude-swift is absent by design and fails closed. The version here is @ant/desktop v1.24012.9; a different version will not necessarily match the env-gated patches, which are keyed to strings in this build. npm install stock Electron 42.9.2, once ./run.sh Electron 42.9.2 is pinned to match the native modules harvested from a current Claude Desktop install. A window opens on the claude.ai login screen; sign in as usual. run.sh re-creates the resources/app.asar → app/ symlink and reinstalls the disclaimer helper on every launch, so reinstalling node modules costs nothing. Linux note: npm install downloads the Electron binary from GitHub's release CDN. If that host is blocked on your network, use a mirror: ELECTRON MIRROR=https://registry.npmmirror.com/-/binary/electron/ \ ELECTRON CUSTOM DIR="v{{ version }}" node node modules/electron/install.js run.sh is cross-platform uname -s : it resolves the Electron resources dir dist/resources on Linux, Electron.app/Contents/Resources on macOS , the Claude Desktop profile ~/.config/Claude vs ~/Library/Application Support/Claude , and the disclaimer helper path per platform. See LINUX.md https://github.com/mkornreich/llm-desktop-electron/blob/main/docs/LINUX.md for the port's design and the folded-conditional gotchas. On modern Ubuntu kernel.apparmor restrict unprivileged userns=1 Electron's namespace sandbox is blocked, so it needs a setuid-root chrome-sandbox — which the downloaded Electron isn't. run.sh self-heals this by symlinking to the setuid-root chrome-sandbox the Claude Desktop .deb already installed same Electron build, byte-identical , keeping the renderer sandbox on with no sudo . If no system copy matches, it prints the one-time sudo chown root:root … && sudo chmod 4755 … fix. Stock Electron has no equivalent of Claude Desktop's TCC-attribution helper, and the app invokes it as disclaimer