{"slug": "setting-up-pi-with-deepseek-v4-1-flash-on-openrouter", "title": "Setting Up Pi With DeepSeek V4.1 Flash on OpenRouter", "summary": "A developer's audit of pi's session logs against OpenRouter's generation API found a supposedly cheap DeepSeek V4.1 Flash setup paying about three times what it should, because OpenRouter's default Auto Exacto routing ignored price and sent all 123 requests in two sessions to Parasail, which dropped the prompt cache fully five times and partially seven times. The fix pins routing to Fireworks with `allow_fallbacks: false`, `data_collection: \"deny\"` and `zdr: true` in `~/.pi/agent/models.json`, since the twelve cache-miss turns accounted for half the bill at $0.22 to $0.30 per million input tokens versus $0.006 to $0.007 for cached reads.", "body_md": "### How to Make Videos with Claude Code\n\nClaude Code can't generate video directly. It isn't a video model like Veo or Sora. It writes text and code. You can…\n\nWhen I run out of Codex and Claude Code usage, I switch to DeepSeek V4.1 Flash in [pi](https://github.com/earendil-works/pi) through OpenRouter. In [Inference Is the Last LLM Moat](https://www.vincentschmalbach.com/inference-is-the-last-llm-moat/) I called the providers on OpenRouter a mixed bag. This week I read pi's session logs against OpenRouter's generation API and found my supposedly cheap setup paying about three times what it should. This is the configuration I use now to avoid the most common traps.\n\nPi takes custom models from `~/.pi/agent/models.json`. Log in once with `/login openrouter` inside pi, then add this entry:\n\n```\n{\n  \"providers\": {\n    \"openrouter\": {\n      \"models\": [\n        {\n          \"id\": \"deepseek/deepseek-v4.1-flash\",\n          \"name\": \"DeepSeek V4.1 Flash (Fireworks, ZDR)\",\n          \"reasoning\": true,\n          \"thinkingLevelMap\": { \"off\": null, \"minimal\": \"low\", \"low\": \"low\", \"medium\": \"high\", \"high\": \"high\", \"xhigh\": \"max\", \"max\": \"max\" },\n          \"input\": [\"text\"],\n          \"contextWindow\": 262144,\n          \"maxTokens\": 131072,\n          \"cost\": { \"input\": 0.22, \"output\": 0.66, \"cacheRead\": 0.007, \"cacheWrite\": 0 },\n          \"compat\": {\n            \"thinkingFormat\": \"openrouter\",\n            \"sendSessionAffinityHeaders\": true,\n            \"sessionAffinityFormat\": \"openrouter\",\n            \"openRouterRouting\": {\n              \"only\": [\"fireworks\"],\n              \"allow_fallbacks\": false,\n              \"data_collection\": \"deny\",\n              \"zdr\": true\n            }\n          }\n        }\n      ]\n    }\n  }\n}\n```\n\nAnd in `~/.pi/agent/settings.json`:\n\n```\n{\n  \"defaultProvider\": \"openrouter\",\n  \"defaultModel\": \"deepseek/deepseek-v4.1-flash\",\n  \"defaultThinkingLevel\": \"high\",\n  \"retry\": { \"enabled\": true, \"maxRetries\": 5, \"baseDelayMs\": 2000 }\n}\n```\n\nPi passes the `openRouterRouting` object straight through as OpenRouter's `provider` field, so anything from the [provider routing docs](https://openrouter.ai/docs/guides/routing/provider-selection) works there. The `cost` block only feeds pi's cost display. `contextWindow` controls when pi compacts, which I get to at the end.\n\nDeepSeek V4.1 Flash costs $0.22 to $0.30 per million input tokens on the OpenRouter providers I would consider, and a cached read costs $0.006 or $0.007. Pi sends the whole conversation on every turn, so after a few tool calls almost the entire prompt should be cache reads. At 120k tokens of context a turn costs less than a tenth of a cent when the cache hits and 3 to 4 cents when it misses.\n\nAcross 123 turns in two sessions on one day, the cache dropped fully five times and partially seven times. Those twelve turns were half of the bill. I assumed OpenRouter had switched providers on me. It had not. All 123 requests went to Parasail, on the same endpoint ID, and Parasail dropped the cache on its own. On the older V4 Flash 0731 model it was worse, with hits on 5 of 17 turns.\n\nMy old config listed four zero-data-retention providers and let OpenRouter choose between them. For plain requests OpenRouter load-balances by price. Every pi request includes tools, and for tool-calling requests OpenRouter runs [Auto Exacto](https://openrouter.ai/docs/guides/routing/auto-exacto) by default, a routing step that reorders providers by OpenRouter's own tool-calling benchmarks and ignores price. For this model it picks Parasail, which was the most expensive provider in my pool and had the worst uptime. The docs themselves say Auto Exacto can cause cache misses mid-conversation in agent loops, and that `sort: \"price\"` turns it off.\n\nThe prompt cache lives on the provider's endpoint. OpenRouter's sticky routing keeps a session on one provider, and pi already sends the `x-session-id` header it keys on, so that part works without configuration. With `allow_fallbacks: true`, though, one upstream rate limit is enough to move the session to another provider, which has no cache. In one test DeepInfra returned a 429, OpenRouter fell back to Fireworks, and the whole prompt was billed as fresh input. Pinning one provider and turning fallbacks off means a 429 becomes a retry with backoff instead. That is why the retry count in settings.json goes up from three to five. Waiting half a minute on a retry is cheaper than paying four cents for a fallback turn.\n\nI benchmarked the candidates with a 15k-token prompt over 20 turns, one provider pinned at a time, checking `cached_tokens` in every response.\n\n| Provider | Cache hits | Errors | Input / output / cache read per million | \n|---|---|---|---|\n| Fireworks | 19 of 19 | none | $0.22 / $0.66 / $0.007 | \n| Novita | 18 of 19 | none | $0.30 / $1.20 / $0.006 | \n| Parasail | 8 of 9 in a shorter run, worse in real sessions | none | $0.30 / $1.20 / $0.006 | \n| Morph | 9 of 11, with 3 to 19 second latencies | none | $0.195 / $0.78 / $0.006 | \n| DeepInfra | 1 answered out of 12 | 11 rate limits | $0.20 / $0.60 / $0.006 | \n\nCache read pricing matters less than I thought. With the cache holding, cached reads were about a third of my daily bill, and fresh input and output were the other two thirds, so Fireworks at $0.007 per cached million still comes out 20 percent cheaper than Novita overall. The outliers are what to watch for. BaseTen charges $0.03 per million cached tokens for the same model, five times the others, so check all three prices on the model's provider list before pinning anything.\n\nThe rate limits are upstream, per model, shared across everyone on OpenRouter, and they come and go. DeepInfra was unusable for V4.1 Flash but fine for 0731, where it is also the cheapest usable option at $0.06 per million input. Fireworks returned four 429s in ten calls at one point in the afternoon and none in the 31 calls after that. If Fireworks starts stalling on retries, I move the pin to Novita.\n\n`zdr: true` and `data_collection: \"deny\"` restrict routing to providers that do not retain prompts. That rules out DeepSeek's own endpoint, which would be the cheapest at $0.15 per million input and $0.003 per cached million. I do not want client code ending up in a training set, so I accept that.\n\nPi compacts when the context passes `contextWindow` minus a 16,384-token reserve. DeepSeek V4.1 Flash advertises a million tokens, and if you write 1,048,576 into models.json, pi basically never compacts. Context costs more as it grows, even when every read is cached. A miss at a million tokens costs 30 cents and takes a minute to prefill. A small fast model does not get better from hundreds of thousands of tokens of old tool output. Codex by default runs on a 272k window and compacts near the top of it. I set 262144, which puts pi's threshold at about 245k. My long sessions peak around 120k, so it rarely fires, and it caps the damage when one runs away. When I switch tasks inside a session I run `/compact` by hand instead of carrying the old context along.\n\nGive Vroni a GitHub issue, bug report, spec, or rough idea. It reads the repo, plans the change, writes code, runs checks, and works toward a review-ready pull request.\n\nTake a look at vroni.com", "url": "https://wpnews.pro/news/setting-up-pi-with-deepseek-v4-1-flash-on-openrouter", "canonical_source": "https://www.vincentschmalbach.com/pi-deepseek-v4-1-flash-openrouter-setup/", "published_at": "2026-09-16 16:16:28+00:00", "updated_at": "2026-10-01 06:18:06.317720+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-infrastructure", "ai-agents"], "entities": ["pi", "DeepSeek V4.1 Flash", "OpenRouter", "Parasail", "Fireworks", "Claude Code", "Codex", "Vincent Schmalbach"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/setting-up-pi-with-deepseek-v4-1-flash-on-openrouter", "markdown": "https://wpnews.pro/news/setting-up-pi-with-deepseek-v4-1-flash-on-openrouter.md", "text": "https://wpnews.pro/news/setting-up-pi-with-deepseek-v4-1-flash-on-openrouter.txt", "jsonld": "https://wpnews.pro/news/setting-up-pi-with-deepseek-v4-1-flash-on-openrouter.jsonld"}}