5 October 2026
I have continued to fiddle around with how web search works on my local LLM setup. As I’ve mentioned before, it’s super important not to neglect the search function. Even a smaller model can do great work provided it can look up correct and up-to-date information.
Providers #
For a couple of weeks I used my own port of SearXNG. When you look into this project closely, though, you realise that minor matters like “complying with terms of service” aren’t really a priority. This issue captures the flavour. Many search backends are putting up Anubis or captcha challenges and it’s abundantly clear that they don’t want LLMs autonomously using their search index. And y’know, fair enough. It’s a service that costs money to run and it’s only reasonable that someone pays someone for it, especially if they’re not going to be looking at the ads.
To that end I tried out both Kagi and Brave Search’s APIs. I picked those two simply because I recognised both brands and their commitment to privacy. That’s pretty important if you’re going to funnel all of your search queries through them. In the end, Brave won on price. Today, it’s 0.5c per search vs 1.2c for Kagi. Brave also has a generous 1000/month free tier. I expect I’ll exceed that but paying a few bucks a month for this kind of service seems like a fair deal.
Happily, both Open WebUI and OMP make it straightforward to use. In the former you can choose your provider under Web Search settings and paste in the API key; in the latter you set the envvar BRAVE_API_KEY and it picks it up automatically.
Web search flow and cache reuse #
Now, let’s talk about performance. I had a look at the logs from llama-server while having a conversation through Open WebUI and I noticed a bit of context churn. To save memory, I currently have no RAM cache configured, which means that if the model swaps from one context to another then back again, my GPU has to prefill that entire conversation from the start. Cache can help but it comes with its own trade-offs. In general it’s better to avoid switching contexts in a local inference setup.
For some reason, this switching back-and-forth was happening even during a single linear conversation in OWUI. The impact wasn’t outrageous—chat conversations tend to have short contexts compared with coding sessions so it caught up quickly. Still, it would be nice to avoid those unnecessary seconds of latency if I can. I captured some verbose logs to check what actually happened through a multi-turn conversation:
The first thing to notice is that it wasn’t the web search’s fault. If you haven't seen this before, the way search_web and fetch_url work is via tool calls. The model writes some structured markup saying “please do a search for these terms” and then ends its response.
- OWUI sees this special output in the result.
- OWUI runs the search.
- OWUI appends the results to the end of the context as structured JSON.
- OWUI asks the model to continue the conversation.
This means that when I typed one message and pressed enter, this resulted in two distinct calls to the model with a break in the middle where OWUI had to go and run a web search. In this case the model decided it could answer my question from the little snippets in the search results alone, so it gave an answer and stopped. Fetching the content of a URL works exactly the same way.
Both of these are the happy path—OWUI is just throwing more stuff on the end, making the context bigger. Since everything before that is the same, llama-server recognises that it has most of this conversation in cache already and it only has to prefill the new part on the end. Life is good.
However, there are three situations in this log where the cache was not fully engaged, and these have varying levels of impact.
Title generation #
The first request against the model is not even my question. OWUI is asking it to summarise the query in a few words to generate a title. This is pretty useful for browsing conversations later but it adds about 2 seconds latency before it gets to actually answering.
### Task:
Generate a concise title summarizing the chat history.
### Guidelines:
- The title should clearly represent the main theme or subject of the conversation.
...
Open WebUI actually sends two requests in parallel to generate the title and answer the question. The title query comes slightly earlier. Since the GPU can only handle one request at a time, answering the question gets queued behind. Since the prompt for each of these requests is different, this naturally clears the KV cache and we start from scratch. This isn’t back-and-forth churn exactly, but it’s mildly inelegant that both requests contained your query text and the GPU had to prefill it twice.
You can choose under Interface settings whether you want this feature. I’ve decided to leave it on.
Tag generation #
This is the painful one. After the model completed its first reply in full—the combination of two responses and one set of search results—OWUI sent in a separate prompt asking for some “tags” to match this conversation.
### Task:
Generate 1-3 broad tags categorizing the main themes of the chat history, along with 1-3 more specific subtopic tags.
### Guidelines:
- Start with high-level domains (e.g. Science, Technology, Philosophy, Arts, Politics, Business, Health, Sports, Entertainment, Education)
- Consider including relevant subfields/subdomains if they are strongly represented throughout the conversation
...
The model dutifully replied with the list “General”, ejecting the main conversation context in progress.
Then when I typed in a follow-up question it had to eject its cache again and prefill 5,716 tokens from the beginning, which added 4.1 seconds to my next query.
I don’t use tags, so I went under Interface and turned this feature off.
Citation instructions #
OWUI did something quirky here—after the model asked to fetch the full content of a page, it actually edited the human’s third turn. It prepended some special system instructions explaining how to provide inline citations to URLs that it’s referencing.
### Task:
Respond to the user query using the provided context, incorporating inline citations in the format [id] **only when the <source> tag includes an explicit id attribute** (e.g., <source id="1">).
### Guidelines:
- If you don't know the answer, clearly state that.
- If uncertain, ask the user for clarification.
...
This is a useful feature but because it modifies history, it doesn’t match llama-server’s cache. Fortunately, it can recognise that it’s only a couple of hundred tokens on the end which are different. As a result, the practical impact is minimal, and the extra instructions provide important functional value.
Conclusion #
I’m now satisfied I know exactly what’s going on, and with tag generation disabled I’m not seeing any major cache misses. It goes to show how much these front ends are primarily designed for cloud inference where the cost of starting a second context in parallel is minimal. Those of us with only one slot need to be vigilant for any thrashing, or we get stuck with slower speeds.
Serious Computer Business Blog by Thomas Karpiniec