{"slug": "show-hn-fast-ollama-like-router-that-gives-a-text-only-model-eyes-and-ears", "title": "Show HN: Fast Ollama like router that gives a text-only model eyes and ears", "summary": "A developer published funcroute, a roughly 4,000-line C router that sits between a client and model vendors and routes each request to a specialist upstream based on its attachments, letting a cheap text-only model handle images, audio and PDFs. The tool sorts content parts by kind (image_url, input_audio, file), picks a provider ordered audio > file > image, rewrites the model field, and translates Ollama-style replies back to OpenAI dialect, trimming conversations to a configured budget without dropping attachment-bearing messages. It is a personal tool, not an inference engine, load balancer or weight merger, and ships as five shared libraries with no vendored code.", "body_md": "A small C router that gives a text-only model eyes, ears and a filing cabinet.\n\nPoint your client at funcroute instead of at a model vendor, and the cheap text-only model you already use starts handling screenshots, voice memos and PDFs. Your client keeps sending one URL and one model name. It never finds out that a different model answered.\n\n```\nyour client ──▶ funcroute ──┬──▶ deepseek-v4-flash           (text)\n  one URL,                  ├──▶ deepseek/deepseek-v4.1-flash (images, PDFs)\n  one model name            └──▶ qwen/qwen3.8-omni-flash      (audio)\n```\n\nThis is a personal tool that turned out to be reliable enough to publish. It is\nnot a model, not a fine-tune, and not a merger. It picks an upstream per request\nby looking at what is attached to the request, rewrites the `model` field,\nforwards it, and translates the answer back. Roughly 4,000 lines of C, five\nshared libraries, no vendored code.\n\n**What it isn't:** an inference engine, a proxy for your whole stack, a\nload balancer, or a way to make an expensive model cheap. It does not merge\nweights, change prompts, or reformat your messages beyond the model field and\nthe dialect translation. See [Prior art](#prior-art-and-why-this-exists-anyway)\nfor the grown-up tools that do more.\n\nContents: [Measured numbers](#measured-numbers) ·\n[Prior art](#prior-art-and-why-this-exists-anyway) ·\n[Build and run](#build-and-run) · [Configuration](#configuration) ·\n[Reproducing these numbers](#reproducing-these-numbers) ·\n[Sharp edges](#sharp-edges)\n\nMultimodal is a configuration tax. Everyone wants it, nobody wants to pay for it, and every workaround has a taste:\n\n| What I tried | What happened | \n|---|---|\n| Point the client at a frontier omni model | Works perfectly. The bill has a comma in it. | \n| Add a second \"vision\" provider to the client | The client sends one model name, not a routing table. Partway through a conversation it switches back to the text model and 400s on the image still sitting in the history. | \n| Use the vision model for everything | Now I am paying vision prices to reformat JSON, and I lost my cheap prefix cache. | \n| Use the vendor's own multimodal endpoint | The attachment is `image_url` here, base64`images[]` there,`input_audio` somewhere else and`file_data` somewhere else again. Every API has its own dialect for \"here is a picture\". | \n| Hardcode the routing into the application | The application is now a router, and I am now the person who maintains a router. | \n\nUnderneath all of that is a boring fact: models are specialists. Text models can't see, vision models can't hear, and the audio model is not going to read your PDF. \"Multimodal\" is really \"three models and a switchboard\", and most people end up writing the switchboard badly, once per project. This is the switchboard, written once, in C, with no dependencies you don't already have.\n\n1. Walks the message content parts and sorts the attachments by kind:\n`image_url` ,`input_audio` ,`file` . The Ollama`images[]` array counts as an\nimage part, because that is all it is.\n2. Picks a provider for that one request. No attachment means the cheap text model. Attachments mean the specialist for the most restrictive kind in the request, ordered audio > file > image, on the grounds that whatever gets picked has to be able to read everything in the body. The omni model reads images and PDFs too, so it can take an image and a voice memo at once.\n3. Rewrites the `model` field, forwards the request, and translates the reply\nback if the client is speaking Ollama instead of OpenAI.\n4. Trims the conversation to fit the configured budget, without ever dropping a message that carries an attachment.\n\nEverything else passes through untouched. Tool messages are dropped atomically\nwhen trimming, so an assistant `tool_calls` message never survives without its\nrun of `tool` results.\n\nI generated a PNG with a red background, recorded a 440 Hz WAV, and put a secret code in a PDF, then pointed each candidate model at all three. Results, not marketing:\n\n| Model | text | image | audio | file | \n|---|---|---|---|---|\n| `deepseek-v4-flash` (default route) | yes | no | no | no | \n| `deepseek/deepseek-v4.1-flash` (DeepSeek VL) | yes | yes | no | yes | \n| `qwen/qwen3.8-omni-flash` | yes | yes | yes | yes | \n\nThrough the router, on one machine, in one run, one request per row:\n\n| Request | Routed to | Body in | Body out | Wall clock | \n|---|---|---|---|---|\n| text only | deepseek | 107 B | 589 B | 1014 ms | \n| 4 KB image | deepseek-vl | 434 B | 2482 B | 1237 ms | \n| 600 B PDF | deepseek-vl | 1087 B | 1581 B | 2028 ms | \n| 1 s WAV | qwen-omni | 42 953 B | 1879 B | 2471 ms | \n| image + WAV | qwen-omni | 43 306 B | 1718 B | 2992 ms | \n| image, streamed | deepseek-vl | 451 B | 63 902 B | 1909 ms | \n| image via `/api/chat` | deepseek-vl | 326 B | 625 B | 703 ms | \n| image via `/api/generate` | deepseek-vl | 293 B | 664 B | 738 ms | \n\nThe useful part of that table is the routing column, not the latency. Every audio-bearing row went to a different vendor than every text-only row, and the client asked for the same model name throughout.\n\nTwo findings shaped the shipped config, both from that test run.\n\n**DeepSeek VL has no ears.** No DeepSeek endpoint takes `input_audio` at all.\nOpenRouter answers `404 no endpoints found that support input audio` and the\nfirst-party API rejects the part type outright. Audio needed its own route to an\nomni model, which is why routing is per kind rather than one \"media provider\"\nfield.\n\n**DeepSeek VL thinks before it sees.** It is a reasoning model, and with\n`reasoning: high`, or with the field left unset, it can spend a small client\n`max_tokens` entirely on its reasoning trace and hand back empty content. In the\ntest that was 0 to 2 successes out of 3 depending on the run. At `reasoning: low`\nit answered correctly 3 out of 3. Both attachment providers ship at `low`.\n\n| Request contains | Goes to | Model | \n|---|---|---|\n| nothing special | DeepSeek | `deepseek-v4-flash` | \n| an `image_url` part | OpenRouter | `deepseek/deepseek-v4.1-flash` | \n| a `file` part | OpenRouter | `deepseek/deepseek-v4.1-flash` | \n| an `input_audio` part | OpenRouter | `qwen/qwen3.8-omni-flash` | \n\nImages and PDFs go through OpenRouter's DeepSeek VL rather than the first-party\nDeepSeek API because `api.deepseek.com` wants a pre-uploaded `file_id` before it\nwill look at a file, and rejects inline base64. Keeping both kinds on the same\nupstream also means a mixed request stays with one provider.\n\nAttachments travel inline in the request body as base64 data URLs. funcroute never opens a file and accepts no multipart uploads. That is on purpose: the things it routes are already coming off a stream somewhere, so the client has the bytes before it has a path. If you want a PDF to reach the router by name, that is your client's problem, and it is four lines of base64 in any language.\n\nThere are grown-up projects in this space and you should look at them first. LiteLLM and Portkey both do provider abstraction with far more, Bifrost and the Kong, Envoy and Cloudflare gateways are built for teams, and OpenRouter is already an aggregator with its own routing.\n\nThe difference that matters here is *when* the decision happens. Those tools\ngenerally pick an upstream from the model name you asked for, or fall back when\none errors or is rate limited. funcroute ignores the model name on the way in\n(it rewrites it) and picks from the content parts in the request body. That is\nthe whole reason it can sit behind a client that has one model name hardcoded\nand no idea what a vision model is.\n\nThings that follow from that being the only goal:\n\n- It is C with five `pkg-config` dependencies, so the binary is one file with no\nruntime. It starts in milliseconds and you can read all of it in a sitting.\n- It is not a service, a dashboard, or a config DSL. There is one JSON config.\n- It does not do prompt rewriting, semantic routing, caching, budgets, retries,\nor load balancing across keys. If you want those, use something else, and\nconsider running it *behind* one of them.\n- It speaks both the OpenAI and Ollama dialects, because the two clients I actually use do not agree on which one is correct.\n\n```\nmake                      # libcurl, jansson, libmicrohttpd, openssl, sqlite3\ncp .env.example .env      # then put DEEPSEEK_API_KEY / OPENROUTER_API_KEY in it\nchmod 600 .env\n./run.sh\n```\n\n`run.sh` sources `.env` and execs the binary. You should see this:\n\n```\nfuncroute 0.5.0: listening on 127.0.0.1:11434 (endpoint /v1/chat/completions)\n  text  -> provider \"deepseek\"    model \"deepseek-v4-flash\"\n  image -> provider \"deepseek-vl\" model \"deepseek/deepseek-v4.1-flash\"\n  file  -> provider \"deepseek-vl\" model \"deepseek/deepseek-v4.1-flash\"\n  audio -> provider \"qwen-omni\"   model \"qwen/qwen3.8-omni-flash\"\n  attachment part types: image \"image_url\", audio \"input_audio\", file \"file\"\n  ollama /api/tags exports model \"bhag\"\n```\n\nFour routes, one port. Point anything OpenAI-shaped or Ollama-shaped at\n`http://127.0.0.1:11434` and ask it about a picture.\n\n```\n# text, handled by DeepSeek\ncurl -s localhost:11434/v1/chat/completions -H 'Content-Type: application/json' \\\n  -d '{\"model\":\"bhag\",\"messages\":[{\"role\":\"user\",\"content\":\"hello\"}]}'\n\n# image, handled by DeepSeek VL\ncurl -s localhost:11434/v1/chat/completions -H 'Content-Type: application/json' \\\n  -d '{\"model\":\"bhag\",\"messages\":[{\"role\":\"user\",\"content\":[\n        {\"type\":\"text\",\"text\":\"what is this?\"},\n        {\"type\":\"image_url\",\"image_url\":{\"url\":\"data:image/png;base64,<...>\"}}]}]}'\n\n# audio, handled by the omni model\ncurl -s localhost:11434/v1/chat/completions -H 'Content-Type: application/json' \\\n  -d '{\"model\":\"bhag\",\"messages\":[{\"role\":\"user\",\"content\":[\n        {\"type\":\"text\",\"text\":\"transcribe this\"},\n        {\"type\":\"input_audio\",\"input_audio\":{\"data\":\"<base64 wav>\",\"format\":\"wav\"}}]}]}'\n\n# a PDF, handled by DeepSeek VL\ncurl -s localhost:11434/v1/chat/completions -H 'Content-Type: application/json' \\\n  -d '{\"model\":\"bhag\",\"messages\":[{\"role\":\"user\",\"content\":[\n        {\"type\":\"text\",\"text\":\"summarise this\"},\n        {\"type\":\"file\",\"file\":{\"filename\":\"report.pdf\",\n         \"file_data\":\"data:application/pdf;base64,<...>\"}}]}]}'\n```\n\nSame URL, same model name, four different upstream models.\n\nAll external, all via `pkg-config`, nothing hand-rolled:\n\n| Library | What for | \n|---|---|\n| libmicrohttpd | the HTTP server | \n| libcurl | talking to the upstreams | \n| jansson | reading and rewriting JSON | \n| SQLite | the optional request log | \n| OpenSSL (EVP) | base64 | \n| pthreads | the SSE producer thread | \n\nBuilds clean with `-std=c2x -Wall -Wextra -Wpedantic`, no warnings.\n\nTagged builds are attached to the releases page, one tarball per platform:\n\n| Tarball | Built on | Notes | \n|---|---|---|\n| `funcroute-<version>-linux-x86_64.tar.gz` | ubuntu-24.04 | links your distro's libcurl, jansson, libmicrohttpd, openssl, sqlite3 | \n| `funcroute-<version>-linux-aarch64.tar.gz` | ubuntu-24.04-arm | same | \n| `funcroute-<version>-darwin-arm64.tar.gz` | macos-15 | self-contained, the Homebrew dylibs ship in `lib/` | \n| `funcroute-<version>-darwin-x86_64.tar.gz` | macos-15-intel | same | \n\nEach tarball unpacks into a single directory holding `funcroute`,\n`funcroute-client`, `run.sh`, `config.json`, `.env.example`, the README and the\nlicence, so `./run.sh` works straight out of the unpack. `SHA256SUMS` covers all\nof them.\n\nAll four targets are built on native runners, so nothing is cross-compiled or\nemulated and no emulation shows up in your timings. The macOS tarballs are the\nonly ones doing real work: `scripts/package.sh` walks `otool -L`, copies every\nnon-system dylib into `lib/`, rewrites each reference to `@executable_path/lib/`,\nand then fails the build if anything still points at the Homebrew prefix. The\nLinux binaries are plain dynamic executables, which is why they are a quarter of\na megabyte instead of 40 MB, and why your distro needs the libraries:\n\n```\napt-get install libcurl4 libjansson4 libmicrohttpd12 libssl3 libsqlite3-0  # Debian/Ubuntu\ndnf install libcurl jansson libmicrohttpd openssl sqlite                   # Fedora\npacman -S curl jansson libmicrohttpd openssl sqlite                        # Arch\n```\n\nOr the old way, from source:\n\n```\nmake VERSION=0.5.0\nsudo make install              # /usr/local/bin; override with PREFIX=/usr\n```\n\nThe version is compiled into both binaries. It shows up in the startup banner\nand in `GET /api/version`, so you can tell a release build from a `git describe`\nbuild without guessing. A binary compiled by hand, without `VERSION`, reports\n`dev`.\n\nCI runs on every push and pull request: build, `make test`, package, on all four\nplatforms. That test is offline (mock upstreams, no keys, no network) and asserts\nthe routed provider and model recorded for each attachment kind, so a routing\nregression fails the build instead of a release. Cutting a release is\n`git tag v0.5.1 && git push origin v0.5.1`; the workflow builds, tests, packages,\nwrites checksums and publishes. Re-running it on the same tag replaces the\nassets.\n\nJSON, read from `config.json` unless you pass a path as the first argument or\nset `FUNCROUTE_CONFIG`.\n\n```\n{\n  \"server\":     { \"host\": \"127.0.0.1\", \"port\": 11434, \"max_connections\": 128 },\n  \"database\":   { \"path\": \"funcroute.db\" },   // \"\" disables persistence\n\n  \"providers\": {\n    \"deepseek\": {\n      \"type\": \"openai\",                        // reserved for future use\n      \"base_url\": \"https://api.deepseek.com\",  // the endpoint path is appended\n      \"api_key_env\": \"DEEPSEEK_API_KEY\",       // or \"api_key\": \"<literal>\"\n      \"model\": \"deepseek-v4-flash\",\n      \"reasoning\": \"high\",                     // becomes reasoning_effort\n      \"timeout_secs\": 120\n    },\n    \"deepseek-vl\": {\n      \"type\": \"openai\",\n      \"base_url\": \"https://openrouter.ai/api\",\n      \"api_key_env\": \"OPENROUTER_API_KEY\",\n      \"model\": \"deepseek/deepseek-v4.1-flash\", // eyes: images and PDFs\n      \"reasoning\": \"low\",                      // low is what keeps content non-empty\n      \"timeout_secs\": 180\n    },\n    \"qwen-omni\": {\n      \"type\": \"openai\",\n      \"base_url\": \"https://openrouter.ai/api\",\n      \"api_key_env\": \"OPENROUTER_API_KEY\",\n      \"model\": \"qwen/qwen3.8-omni-flash\",      // the only route with ears\n      \"reasoning\": \"low\",\n      \"timeout_secs\": 180\n    }\n  },\n\n  \"routing\": {\n    \"default_provider\": \"deepseek\",      // text-only requests\n    \"image_provider\": \"deepseek-vl\",     // image_url parts\n    \"file_provider\": \"deepseek-vl\",      // file parts\n    \"audio_provider\": \"qwen-omni\",       // input_audio parts\n    \"media_provider\": \"deepseek-vl\",     // catch-all; empty means image_provider\n\n    \"image_content_type\": \"image_url\",   // which part types count as what\n    \"audio_content_type\": \"input_audio\",\n    \"file_content_type\": \"file\",\n\n    \"endpoint\": \"/v1/chat/completions\",\n    \"ollama_model\": \"bhag\",              // the name /api/tags hands out\n    \"max_request_bytes\": 131072,         // context budget, 0 turns it off\n    \"max_messages\": 128,\n    \"advertised_context_length\": 1000000\n  }\n}\n```\n\nThe attachment routes fall back to each other, specific beats general. Leave\n`media_provider` out and it inherits `image_provider`. Leave `file_provider` or\n`audio_provider` out and they inherit `media_provider`. Old configs that only\nknew about `image_provider` and `media_provider` still start; they just don't get\nper-kind routing. If a route names a provider that doesn't exist, that is a\nstartup error and not a surprise at three in the morning.\n\nClients resend the entire conversation on every turn, and screenshots and PDFs are megabytes. Without a cap, one pasted image pushes the whole history out the window.\n\n- `max_request_bytes` and`max_messages` bound what gets forwarded. The router\ndrops the oldest messages until both fit.\n- The system prompt stays first and byte-for-byte identical, so DeepSeek's automatic prefix caching keeps hitting the part of the prefix that survived.\n- The bytes inside an attachment are not counted, and a message carrying an attachment is never dropped. Otherwise the base64 for one screenshot would look enormous and take the screenshot with it.\n- A tool exchange goes together. If an assistant `tool_calls` message is\ndropped, its run of`tool` results goes with it. An orphaned`tool` message\nwith no predecessor is dropped regardless.\n- All of this applies to `/v1/chat/completions` ,`/api/chat` and`/api/generate` .\n\n`advertised_context_length` is the number clients show in their context meter,\nserved through `/v1/models`, `/api/tags` and `/api/show`. It is not a promise\nabout what actually gets forwarded, since `max_request_bytes` decides that. Set\nit high enough that your client doesn't start panicking and trimming before the\nrouter's own budget has had a chance to do anything.\n\n```\n./run.sh                 # source .env, then exec ./funcroute\n./run.sh config.json     # explicit config\nENV_FILE=prod.env ./run.sh\nLOGDB=off ./run.sh       # or: LOGDB=/var/tmp/funcroute.db\n```\n\n`.env` is plain `KEY=VALUE` and is sourced with `set -a`, so you don't need\n`export` in it. A missing `.env` is a warning, not a failure. If you would rather\nrun the binary directly, export `DEEPSEEK_API_KEY` and `OPENROUTER_API_KEY`\nyourself and skip the script.\n\nSQLite persistence is opt in per run. Command line beats environment beats\n`config.json`.\n\n```\n./funcroute                       # log to database.path\n./funcroute --log other.db        # --db is an alias\n./funcroute --no-log              # no database opened or created at all\nFUNCROUTE_LOG=other.db ./funcroute\nFUNCROUTE_NO_LOG=1 ./funcroute\n```\n\nSetting `\"database\": {\"path\": \"\"}` turns it off too. Disabled means nothing is\nopened and nothing is created, and you still get the one line summaries on\nstdout, which are the useful part:\n\n``` php\n[2026-10-02T02:56:29Z] openai /v1/chat/completions model=bhag\n  -> deepseek-vl/deepseek/deepseek-v4.1-flash (image) req=434B resp=2482B status=200 dur=1237ms\n```\n\nThose `(image)`, `(audio)` and `(file)` markers are there so you can convince\nyourself the picture really did go somewhere other than where the text went.\n\nThe `requests` table holds `id, ts, protocol, endpoint, requested_model, routed_provider, routed_model, has_image, has_audio, has_file, stream, request_size, response_size, status, duration_ms, trimmed_bytes, error, request_body, response_body`. An older database gets migrated in place with\n`ALTER TABLE ADD COLUMN`, so you don't lose it. Request and response bodies are\nstored in full by default, which is worth knowing before you log audio.\n\nfuncroute speaks Ollama as well, so Ollama-native clients work without changes.\n\n| Route | Method | Notes | \n|---|---|---|\n| `/api/chat` | POST | translates between OpenAI parts and Ollama `images[]` | \n| `/api/generate` | POST | single prompt | \n| `/api/tags` | GET | exports `ollama_model` ,`bhag` by default | \n| `/api/show` | POST | model card plus the advertised context | \n| `/api/version` ,`/api/ps` | GET | stubs, so probing clients calm down | \n| `/api/embeddings` | POST | answers `501 not supported` , honestly | \n\nStreaming is converted in both directions on a producer thread, upstream SSE on\none side and Ollama NDJSON on the other. Reasoning traces are re-surfaced rather\nthan dropped: DeepSeek's `reasoning_content` and OpenRouter's `reasoning` both\ncome out as Ollama `message.reasoning`, or a top level `reasoning` for\n`/api/generate`. Thinking models keep thinking through the router.\n\nThe caveat is that Ollama's API can only express images. Audio and file parts\nhave nowhere to go, so `funcroute-client` refuses them in `--ollama` mode instead\nof quietly throwing your attachment away.\n\n`funcroute-client` is a dependency-free CLI:\n\n| Flag | Meaning | \n|---|---|\n| `-c/--config` ,`-b/--base-url` ,`-e/--endpoint` ,`-m/--model` | where to talk | \n| `-p/--prompt` ,`-f/--file` | what to send | \n| `-i/--image URL\\|PATH` | attach an image, local files become data URLs | \n| `-a/--audio PATH` ,`-A/--attach PATH` | attach audio, or a file such as a PDF | \n| `-s/--stream` | stream tokens as SSE | \n| `-o/--ollama` | use `/api/chat` instead | \n| `-t/--timeout` ,`-h/--help` | the usual | \n\n```\n./funcroute-client -i screenshot.png -p \"what broke?\"\n./funcroute-client -a memo.wav -p \"summarise this\"\n./funcroute-client -A report.pdf -p \"what is the account number?\"\n./funcroute-client -i chart.png -a note.wav -p \"describe both\"   # lands on omni\n```\n\n`frontend.html` is the one I actually use for poking at it: open it in a browser,\npoint it at the router, drop in images, audio or PDFs. It classifies each file by\nkind, builds the right content part, and shows each one as a chip you can remove\nbefore sending.\n\nEverything in the tables above is a script in this repo, so you can check it rather than trust it. The fixtures are generated byte by byte (no downloads) and each one has a checkable answer in it: a red square, a 440 Hz tone, and one secret code.\n\n```\npython3 test/bench/gen_fixtures.py test/bench/fixtures   # PNG, WAV, PDF + payloads\n\n# which models accept which attachment kind (the capability table)\nDEEPSEEK_API_KEY=... OPENROUTER_API_KEY=... \\\n  python3 test/bench/probe_capabilities.py\n\n# reasoning effort vs empty content (the reasoning: low finding)\nOPENROUTER_API_KEY=... python3 test/bench/probe_reasoning.py\n\n# end to end through the router: every row of the measured table\npython3 test/bench/e2e_live.py\n\n# offline routing test: mock upstreams, no keys. This is the CI gate.\nmake test\n```\n\n`e2e_live.py` starts the router on a scratch database, sends one request per\nattachment kind, streams one of them, exercises both Ollama routes, prints the\nlogged routing decision for each request, and exits non-zero if anything failed.\nIt needs real keys and costs a few cents. `probe_capabilities.py` prints the raw\nstatus code and the model's own reply per cell, so a `404` there is evidence,\nnot my summary of one.\n\nFor routing decisions without spending anything, there are mock upstreams:\n\n```\npython3 test/mock_upstream.py 9101 deepseek &\npython3 test/mock_upstream.py 9102 openrouter &\n./funcroute test/config.test.json --no-log\n```\n\nThe mocks echo which provider handled the request and emit provider-native\nreasoning fields. `test/config.test.json` mirrors the production routing on port\n11434, per-kind routes included.\n\n```\nsrc/server.c        HTTP server, routing decision, context trimming   (~980 lines)\nsrc/ollama.c        Ollama dialect, NDJSON streaming, format conversion (~1330 lines)\nsrc/client.c        the CLI client                                     (~840 lines)\nsrc/config.c/.h     JSON config, provider resolution, fallbacks        (~360 lines)\nsrc/provider.c/.h   one upstream request, reasoning field handling     (~175 lines)\nsrc/logdb.c/.h      optional SQLite request log                        (~285 lines)\ntest/routing_test.py    offline routing assertions against mock upstreams\ntest/mock_upstream.py   mock OpenAI upstream for routing tests\ntest/bench/*.py         the scripts behind the measured numbers\nscripts/package.sh      builds a release tarball for one platform\n.github/workflows/      the 4-platform build matrix and the tag-driven release\n```\n\nAbout 4,000 lines of C in total, of which roughly 3,200 is the router and 840 is the client.\n\n| Route | Method | Notes | \n|---|---|---|\n| `/v1/chat/completions` | POST | OpenAI chat, streaming or not | \n| `/v1/models` | GET | OpenAI-style model list | \n| `/api/chat` ,`/api/generate` | POST | Ollama dialect | \n| `/api/tags` ,`/api/show` ,`/api/version` ,`/api/ps` ,`/api/embeddings` | varies | the Ollama probing surface | \n\nThe honest list, in the order you would hit them:\n\n- **Audio has exactly one route.** DeepSeek has no audio input at all, so if`qwen-omni` is down, audio is down. Images and PDFs carry on without it.\n- **Attachment routes must stay on `reasoning: low`.** A reasoning model can eat\na small`max_tokens` and return nothing, which looks exactly like a broken\nrouter. Measured above.\n- **Routing is per request, not per conversation.** A conversation that starts\nas text and then pastes a screenshot will be answered by two different models.\nEach request is stateless, so the vision model sees the whole history it needs,\nbut do not expect a stable \"the model\" across a session.\n- **Bodies are logged in full by default.** A 43 KB audio request becomes a 43 KB\ndatabase row. Disable the log or trim it if that bothers you.\n- **Two upstreams are paid third parties.** The keys go in`.env` , which is`chmod 600` and gitignored. Your attachments leave your machine, obviously.\n- **Prefix caching only survives where the prefix survives.** Trimming keeps the\nsystem prompt byte-identical for exactly that reason, so don't reorder messages\nupstream of the router and expect the cache to follow along.\n- **The test suite is thin.**`make test` runs`test/routing_test.py` , which\nstarts mock upstreams and asserts the routed provider and model recorded for\nevery attachment kind. That runs in CI on all four platforms with no keys. The\nlive probes in`test/bench` are not run in CI, because they need real accounts\nand cost money; run them yourself if you doubt a number above.\n- **This is capability stitching, not a model merger.** The composite only looks\nlike one big multimodal model because the router swaps upstreams per request.\nPlease don't cite it in a paper.\n- **The economics are modest.** Cheapskate the text onto the cheap model and let\nonly attachments reach the expensive one. It doesn't make the expensive model\ncheap, it makes it rare.\n\nMIT. See [LICENSE](https://github.com/SpaceSwordAI/funcroute/blob/main/LICENSE). Use it, fork it, ship it.", "url": "https://wpnews.pro/news/show-hn-fast-ollama-like-router-that-gives-a-text-only-model-eyes-and-ears", "canonical_source": "https://github.com/SpaceSwordAI/funcroute", "published_at": "2026-10-02 03:17:10+00:00", "updated_at": "2026-10-02 03:45:36.006411+00:00", "lang": "en", "topics": ["ai-tools", "large-language-models", "ai-infrastructure", "developer-tools"], "entities": ["funcroute", "Ollama", "OpenAI", "deepseek-v4-flash", "deepseek/deepseek-v4.1-flash", "qwen/qwen3.8-omni-flash"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/show-hn-fast-ollama-like-router-that-gives-a-text-only-model-eyes-and-ears", "markdown": "https://wpnews.pro/news/show-hn-fast-ollama-like-router-that-gives-a-text-only-model-eyes-and-ears.md", "text": "https://wpnews.pro/news/show-hn-fast-ollama-like-router-that-gives-a-text-only-model-eyes-and-ears.txt", "jsonld": "https://wpnews.pro/news/show-hn-fast-ollama-like-router-that-gives-a-text-only-model-eyes-and-ears.jsonld"}}