Why we removed Ollama from Return (and what replaced it) Return, a document analysis tool for lawyers, replaced Ollama with a bundled llama-server from llama.cpp as a sidecar process after Ollama 0.12 introduced cloud models in September 2025 that proxy requests from localhost:11434 to Ollama's datacenter infrastructure. Return's team said the change meant its "Nothing leaked" guarantee rested on a flag rather than a checkable property, since inference locality now depended on daemon configuration, model names, and account tokens outside Return's binary. The replacement llama-server runs with no accounts or hosted model catalogue, binds to 127.0.0.1 on an OS-assigned port, and loads a GGUF file already on disk, with model downloads handled by Return's own code behind an explicit click. Why we removed Ollama from Return and what replaced it Return is a document analysis tool built for lawyers. The people it serves work under privilege, and the first question they ask about an AI tool is always the same: where does the document go? There are two ways to answer. The first is “there’s a setting that keeps everything local, and we keep it set.” The second is “the engine analyzing your document has no way to send it anywhere.” To a security reviewer these are different kinds of claim. This post is about moving from the first to the second, and what it cost. What changed in Ollama 0.12 Until mid-2026, Return’s free tier ran on Ollama. It was a good choice for a long time: a well-maintained local inference server, a clean HTTP API on localhost, a model catalogue our users could draw from. In September 2025, Ollama 0.12 introduced cloud models. The design is clean. A model with a -cloud suffix behaves like any other model: you can list it, run it, point your existing tools at it. Your application keeps sending requests to localhost:11434 exactly as before. The daemon detects the suffix, attaches your account credentials, and proxies the request to Ollama’s datacenter infrastructure. From the client’s perspective, nothing changes. Cloud models are opt-in. They require signing in to an Ollama account, and they only run when someone explicitly selects one. Ollama later added a disable ollama cloud setting and an OLLAMA NO CLOUD environment variable after users asked for a way to enforce local-only operation. Nobody’s document gets uploaded by surprise. Ollama did nothing wrong here, and for plenty of products the hybrid model is exactly right. Why it broke our architecture anyway Our problem was narrower. Return’s promise to lawyers is “Nothing leaked,” and we had been backing it with a checkable statement: run a packet capture while the free tier analyzes a document, and the inference traffic goes to 127.0.0.1 and nowhere else. After 0.12, that statement was still technically true from our process’s point of view. It just no longer meant what it used to. Traffic to localhost:11434 could now be the first hop rather than the destination. Whether inference stayed on the machine had become a function of things outside our binary: which model names existed in the daemon’s list, whether an account token was present, how the daemon was configured, and how that configuration surface might change in future releases. A user on Ollama’s issue tracker put the general worry well when requesting the local-only toggle: even without malice, a misconfiguration or another piece of software could introduce a token or a cloud model without your knowledge. For an individual running Ollama on their own machine, the toggle answers that worry. For a product promising confidentiality to law firms, it doesn’t: the promise now rests on a flag we set rather than on something a reviewer can check. Data protection by design GDPR Article 25 is about what the system can do, not what its settings say today. And a flag hands control of our central claim to a third party’s release schedule. That’s not a criticism of Ollama’s choices. Our requirements are just unusually rigid. From the decision record we wrote at the time: the inability to exfiltrate document content should be a property of the code we ship, not a setting. What we did We replaced Ollama with a bundled llama-server from llama.cpp, run as a sidecar process. llama-server is an inference server and nothing more: no accounts, no hosted model catalogue, no mode in which a request to localhost gets forwarded somewhere else. We start it with -m and the path of a GGUF file that is already on disk. Downloading models is Return’s job, in our own code, behind an explicit click. The sidecar binds to 127.0.0.1 on a port the operating system picks at spawn. Never a fixed well-known port, never 0.0.0.0 : php pub fn pick free port - Result