# Why we removed Ollama from Return (and what replaced it)

> Source: <https://returneditor.ai/blog/not-a-setting-why-we-removed-ollama/>
> Published: 2026-09-28 19:09:22+00:00

# Why we removed Ollama from Return (and what replaced it)

Return is a document analysis tool built for lawyers. The people it serves work under privilege, and the first question they ask about an AI tool is always the same: where does the document go?

There are two ways to answer. The first is “there’s a setting that keeps everything local, and we keep it set.” The second is “the engine analyzing your document has no way to send it anywhere.” To a security reviewer these are different kinds of claim. This post is about moving from the first to the second, and what it cost.

## What changed in Ollama 0.12

Until mid-2026, Return’s free tier ran on Ollama. It was a good choice for a long time: a well-maintained local inference server, a clean HTTP API on localhost, a model catalogue our users could draw from.

In September 2025, Ollama 0.12 introduced cloud models. The design is clean. A model with a `-cloud` suffix behaves like any other model: you can list it, run it, point your existing tools at it. Your application keeps sending requests to `localhost:11434` exactly as before. The daemon detects the suffix, attaches your account credentials, and proxies the request to Ollama’s datacenter infrastructure. From the client’s perspective, nothing changes.

Cloud models are opt-in. They require signing in to an Ollama account, and they only run when someone explicitly selects one. Ollama later added a `disable_ollama_cloud` setting and an `OLLAMA_NO_CLOUD` environment variable after users asked for a way to enforce local-only operation. Nobody’s document gets uploaded by surprise. Ollama did nothing wrong here, and for plenty of products the hybrid model is exactly right.

## Why it broke our architecture anyway

Our problem was narrower. Return’s promise to lawyers is “Nothing leaked,” and we had been backing it with a checkable statement: run a packet capture while the free tier analyzes a document, and the inference traffic goes to `127.0.0.1` and nowhere else.

After 0.12, that statement was still technically true from our process’s point of view. It just no longer meant what it used to. Traffic to `localhost:11434` could now be the first hop rather than the destination. Whether inference stayed on the machine had become a function of things outside our binary: which model names existed in the daemon’s list, whether an account token was present, how the daemon was configured, and how that configuration surface might change in future releases.

A user on Ollama’s issue tracker put the general worry well when requesting the local-only toggle: even without malice, a misconfiguration or another piece of software could introduce a token or a cloud model without your knowledge. For an individual running Ollama on their own machine, the toggle answers that worry. For a product promising confidentiality to law firms, it doesn’t: the promise now rests on a flag we set rather than on something a reviewer can check. Data protection by design (GDPR Article 25) is about what the system can do, not what its settings say today. And a flag hands control of our central claim to a third party’s release schedule. That’s not a criticism of Ollama’s choices. Our requirements are just unusually rigid.

From the decision record we wrote at the time: the inability to exfiltrate document content should be a property of the code we ship, not a setting.

## What we did

We replaced Ollama with a bundled `llama-server` from llama.cpp, run as a sidecar process. `llama-server` is an inference server and nothing more: no accounts, no hosted model catalogue, no mode in which a request to localhost gets forwarded somewhere else. We start it with `-m` and the path of a GGUF file that is already on disk. Downloading models is Return’s job, in our own code, behind an explicit click.

The sidecar binds to `127.0.0.1` on a port the operating system picks at spawn. Never a fixed well-known port, never `0.0.0.0`:

``` php
pub fn pick_free_port() -> Result<u16> {
    let listener = TcpListener::bind(("127.0.0.1", 0))?;
    Ok(listener.local_addr()?.port())
}
```

We bind port zero ourselves, read what the OS assigned, release it, and pass that port to the server explicitly. The original plan was `--port 0` and parsing the resolved port out of the server’s log, but that log line isn’t stable across llama.cpp versions. Pre-binding leaves a small race between releasing the socket and the server taking it; if we lose it, the spawn fails fast and we retry on a fresh port.

Chat and embeddings run as two separate sidecars, because a `llama-server` started in embeddings mode cannot also serve chat, and flipping one process between modes would add model-swap latency to every switch between chatting and searching. Both spawn lazily, so a cold app launch stays cheap.

And we deleted Ollama from the codebase entirely. No feature flag, no dormant `OllamaEngine` kept around just in case. The migration happened before launch, with no production installs to migrate, so the rollback path was `git revert`. Keeping two engine implementations behind a flag would have been permanent debt purchased as insurance against a scenario that version control already covers.

## What it cost us

This trade was not free.

The biggest loss was model management. Ollama is a full application: it pulls models, stores them, lists them, loads them on demand. When we embedded it, all of that lived outside our app, which also meant our users had to install and operate a second program before Return’s free tier did anything. Removing Ollama meant a fresh install of Return has no model at all, and the whole path from zero to working AI became ours to own. We ended up building a models view into Settings: a curated catalogue that ships inside the app and renders offline, plus a Hugging Face browser that reads GGUF metadata over HTTP range requests, so you can check a model’s size and parameters before committing to a multi-gigabyte download. That was engineering time we hadn’t planned to spend. It also produced a better experience than what it replaced: nobody has to install a separate tool or open a terminal to get local AI running.

Model switching became a heavier operation on paper. Loading a different chat model now restarts the whole sidecar rather than swapping weights inside a running daemon, and we budgeted for a visible loading state to cover the gap. In practice the regression never showed up: load time is dominated by reading the weights from disk either way, and on the Apple Silicon hardware we develop on the restart is quick enough that we never built that loading state. Slower machines may feel it more.

Memory went up. A second resident process for embeddings costs roughly 500 to 700 MB with our default model quantized, which lazy spawning mitigates but doesn’t eliminate.

And macOS notarization got harder. Bundling a native binary with its Metal dylibs means every nested Mach-O has to be signed with the hardened runtime before the app bundle is, or notarization rejects it, and the first notarization run with new nested binaries took hours. We now smoke-test that pipeline on a minimal app well before release day.

## Update, September 2026

We no longer ship llama.cpp’s prebuilt binaries. We build the engine ourselves, and we build it without the parts we do not want: the one feature that can run inference on another machine is compiled out, and so is the ability to speak HTTPS. Asked to reach a remote model, the engine simply cannot. On Windows we ship upstream’s build with that same remote-inference part removed.

One caveat, because someone will check: a small model-download helper is still inside the binary. Without HTTPS it cannot fetch anything, and we never give it a URL, but it is there. So the precise claim is narrower than “no networking code”: the engine has no remote inference path, holds no credentials, and is only ever given a local file.

The engine also needs a key that Return generates fresh at every launch, so another program on your computer cannot use it just because it found the port.

## Check it yourself

Return isn’t open source, so “trust us” isn’t an argument. These work on an installed copy:

```
# what the engine links: system libraries and its own dylibs, nothing else
otool -L /Applications/return.app/Contents/MacOS/llama-server

# which llama.cpp build it is (tag and commit are baked in at build time)
/Applications/return.app/Contents/MacOS/llama-server --version

# while a local analysis is running: every engine socket is on 127.0.0.1
lsof -nP -iTCP -a -c llama-server
```

The app itself does use the network: update checks, sign-in and license checks if you have an account, and model downloads from Hugging Face when you ask for one. The claim is about inference. Filter a packet capture to the engine while it analyzes a document and you will see loopback and nothing else.

## What we got

“Where does the document go” is now answered by what we ship rather than by how it’s configured. The free tier is local because the bundled engine has no remote inference path and is only ever given a local model file. That stays true whatever Ollama, or anyone else, ships next quarter. llama.cpp is upstream too, which is why it’s pinned by tag, commit and checksum, and bumped deliberately.

Cloud AI still exists in Return. The paid tier can send requests to a frontier model through our own proxy, under a data processing agreement. That path is a separate, explicit switch in Settings, and the switch says in plain words what leaves the machine. It isn’t a default or a fallback, and configuration drift can’t switch it on. Free is local, and you can check it. Paid is cloud, and you chose it.

Ollama remains good software, and embedding it is the right call for most tools in this space. Our constraint is just unusual: when the premise of a product is that a specific thing cannot happen without the user’s explicit say-so, “cannot” has to live in the architecture. Anything that lives in a setting is a promise. Settings change.
