Show HN: Fast Ollama like router that gives a text-only model eyes and ears A developer published funcroute, a roughly 4,000-line C router that sits between a client and model vendors and routes each request to a specialist upstream based on its attachments, letting a cheap text-only model handle images, audio and PDFs. The tool sorts content parts by kind (image_url, input_audio, file), picks a provider ordered audio > file > image, rewrites the model field, and translates Ollama-style replies back to OpenAI dialect, trimming conversations to a configured budget without dropping attachment-bearing messages. It is a personal tool, not an inference engine, load balancer or weight merger, and ships as five shared libraries with no vendored code. A small C router that gives a text-only model eyes, ears and a filing cabinet. Point your client at funcroute instead of at a model vendor, and the cheap text-only model you already use starts handling screenshots, voice memos and PDFs. Your client keeps sending one URL and one model name. It never finds out that a different model answered. your client ──▶ funcroute ──┬──▶ deepseek-v4-flash text one URL, ├──▶ deepseek/deepseek-v4.1-flash images, PDFs one model name └──▶ qwen/qwen3.8-omni-flash audio This is a personal tool that turned out to be reliable enough to publish. It is not a model, not a fine-tune, and not a merger. It picks an upstream per request by looking at what is attached to the request, rewrites the model field, forwards it, and translates the answer back. Roughly 4,000 lines of C, five shared libraries, no vendored code. What it isn't: an inference engine, a proxy for your whole stack, a load balancer, or a way to make an expensive model cheap. It does not merge weights, change prompts, or reformat your messages beyond the model field and the dialect translation. See Prior art prior-art-and-why-this-exists-anyway for the grown-up tools that do more. Contents: Measured numbers measured-numbers · Prior art prior-art-and-why-this-exists-anyway · Build and run build-and-run · Configuration configuration · Reproducing these numbers reproducing-these-numbers · Sharp edges sharp-edges Multimodal is a configuration tax. Everyone wants it, nobody wants to pay for it, and every workaround has a taste: | What I tried | What happened | |---|---| | Point the client at a frontier omni model | Works perfectly. The bill has a comma in it. | | Add a second "vision" provider to the client | The client sends one model name, not a routing table. Partway through a conversation it switches back to the text model and 400s on the image still sitting in the history. | | Use the vision model for everything | Now I am paying vision prices to reformat JSON, and I lost my cheap prefix cache. | | Use the vendor's own multimodal endpoint | The attachment is image url here, base64 images there, input audio somewhere else and file data somewhere else again. Every API has its own dialect for "here is a picture". | | Hardcode the routing into the application | The application is now a router, and I am now the person who maintains a router. | Underneath all of that is a boring fact: models are specialists. Text models can't see, vision models can't hear, and the audio model is not going to read your PDF. "Multimodal" is really "three models and a switchboard", and most people end up writing the switchboard badly, once per project. This is the switchboard, written once, in C, with no dependencies you don't already have. 1. Walks the message content parts and sorts the attachments by kind: image url , input audio , file . The Ollama images array counts as an image part, because that is all it is. 2. Picks a provider for that one request. No attachment means the cheap text model. Attachments mean the specialist for the most restrictive kind in the request, ordered audio file image, on the grounds that whatever gets picked has to be able to read everything in the body. The omni model reads images and PDFs too, so it can take an image and a voice memo at once. 3. Rewrites the model field, forwards the request, and translates the reply back if the client is speaking Ollama instead of OpenAI. 4. Trims the conversation to fit the configured budget, without ever dropping a message that carries an attachment. Everything else passes through untouched. Tool messages are dropped atomically when trimming, so an assistant tool calls message never survives without its run of tool results. I generated a PNG with a red background, recorded a 440 Hz WAV, and put a secret code in a PDF, then pointed each candidate model at all three. Results, not marketing: | Model | text | image | audio | file | |---|---|---|---|---| | deepseek-v4-flash default route | yes | no | no | no | | deepseek/deepseek-v4.1-flash DeepSeek VL | yes | yes | no | yes | | qwen/qwen3.8-omni-flash | yes | yes | yes | yes | Through the router, on one machine, in one run, one request per row: | Request | Routed to | Body in | Body out | Wall clock | |---|---|---|---|---| | text only | deepseek | 107 B | 589 B | 1014 ms | | 4 KB image | deepseek-vl | 434 B | 2482 B | 1237 ms | | 600 B PDF | deepseek-vl | 1087 B | 1581 B | 2028 ms | | 1 s WAV | qwen-omni | 42 953 B | 1879 B | 2471 ms | | image + WAV | qwen-omni | 43 306 B | 1718 B | 2992 ms | | image, streamed | deepseek-vl | 451 B | 63 902 B | 1909 ms | | image via /api/chat | deepseek-vl | 326 B | 625 B | 703 ms | | image via /api/generate | deepseek-vl | 293 B | 664 B | 738 ms | The useful part of that table is the routing column, not the latency. Every audio-bearing row went to a different vendor than every text-only row, and the client asked for the same model name throughout. Two findings shaped the shipped config, both from that test run. DeepSeek VL has no ears. No DeepSeek endpoint takes input audio at all. OpenRouter answers 404 no endpoints found that support input audio and the first-party API rejects the part type outright. Audio needed its own route to an omni model, which is why routing is per kind rather than one "media provider" field. DeepSeek VL thinks before it sees. It is a reasoning model, and with reasoning: high , or with the field left unset, it can spend a small client max tokens entirely on its reasoning trace and hand back empty content. In the test that was 0 to 2 successes out of 3 depending on the run. At reasoning: low it answered correctly 3 out of 3. Both attachment providers ship at low . | Request contains | Goes to | Model | |---|---|---| | nothing special | DeepSeek | deepseek-v4-flash | | an image url part | OpenRouter | deepseek/deepseek-v4.1-flash | | a file part | OpenRouter | deepseek/deepseek-v4.1-flash | | an input audio part | OpenRouter | qwen/qwen3.8-omni-flash | Images and PDFs go through OpenRouter's DeepSeek VL rather than the first-party DeepSeek API because api.deepseek.com wants a pre-uploaded file id before it will look at a file, and rejects inline base64. Keeping both kinds on the same upstream also means a mixed request stays with one provider. Attachments travel inline in the request body as base64 data URLs. funcroute never opens a file and accepts no multipart uploads. That is on purpose: the things it routes are already coming off a stream somewhere, so the client has the bytes before it has a path. If you want a PDF to reach the router by name, that is your client's problem, and it is four lines of base64 in any language. There are grown-up projects in this space and you should look at them first. LiteLLM and Portkey both do provider abstraction with far more, Bifrost and the Kong, Envoy and Cloudflare gateways are built for teams, and OpenRouter is already an aggregator with its own routing. The difference that matters here is when the decision happens. Those tools generally pick an upstream from the model name you asked for, or fall back when one errors or is rate limited. funcroute ignores the model name on the way in it rewrites it and picks from the content parts in the request body. That is the whole reason it can sit behind a client that has one model name hardcoded and no idea what a vision model is. Things that follow from that being the only goal: - It is C with five pkg-config dependencies, so the binary is one file with no runtime. It starts in milliseconds and you can read all of it in a sitting. - It is not a service, a dashboard, or a config DSL. There is one JSON config. - It does not do prompt rewriting, semantic routing, caching, budgets, retries, or load balancing across keys. If you want those, use something else, and consider running it behind one of them. - It speaks both the OpenAI and Ollama dialects, because the two clients I actually use do not agree on which one is correct. make libcurl, jansson, libmicrohttpd, openssl, sqlite3 cp .env.example .env then put DEEPSEEK API KEY / OPENROUTER API KEY in it chmod 600 .env ./run.sh run.sh sources .env and execs the binary. You should see this: funcroute 0.5.0: listening on 127.0.0.1:11434 endpoint /v1/chat/completions text - provider "deepseek" model "deepseek-v4-flash" image - provider "deepseek-vl" model "deepseek/deepseek-v4.1-flash" file - provider "deepseek-vl" model "deepseek/deepseek-v4.1-flash" audio - provider "qwen-omni" model "qwen/qwen3.8-omni-flash" attachment part types: image "image url", audio "input audio", file "file" ollama /api/tags exports model "bhag" Four routes, one port. Point anything OpenAI-shaped or Ollama-shaped at http://127.0.0.1:11434 and ask it about a picture. text, handled by DeepSeek curl -s localhost:11434/v1/chat/completions -H 'Content-Type: application/json' \ -d '{"model":"bhag","messages": {"role":"user","content":"hello"} }' image, handled by DeepSeek VL curl -s localhost:11434/v1/chat/completions -H 'Content-Type: application/json' \ -d '{"model":"bhag","messages": {"role":"user","content": {"type":"text","text":"what is this?"}, {"type":"image url","image url":{"url":"data:image/png;base64,<... "}} } }' audio, handled by the omni model curl -s localhost:11434/v1/chat/completions -H 'Content-Type: application/json' \ -d '{"model":"bhag","messages": {"role":"user","content": {"type":"text","text":"transcribe this"}, {"type":"input audio","input audio":{"data":"