# Errand Walker: a 1B open model that turns your to-do list into a walk

> Source: <https://dev.to/derbyps/errand-walker-a-1b-open-model-that-turns-your-to-do-list-into-a-walk-40de>
> Published: 2026-10-11 14:53:52+00:00

*This is a submission for the [Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass](https://dev.to/challenges/hacktoberfest-week1-2026-10-05)*

Most short errands happen by default in a car or on a scooter, not because walking is worse, but because planning a walk takes effort. Which shop is closest? In what order? Is it really 40 minutes on foot, or does it just feel like a lot?

**Errand Walker** removes that effort. You type the list the way you'd type it into a notes app:

pharmacy, post office, buy rice

A small open-weight model turns that into structured errands. OpenStreetMap finds the nearest matching places, a routing API computes one loop, and the app answers in a single line: how many stops, how long on foot, how long by car, and whether you can skip the car.

The screen is the shortest part of the experience. Planning takes about six seconds, and then the phone goes back in your pocket. It's a phone-first web app (PWA) for anyone with a handful of small errands in a place with decent OpenStreetMap coverage. It also understands some Indonesian input ("beli obat, kirim paket"), which I evaluated separately from English.

In the video I ask for an ATM, a pharmacy and a meal. The app parses the three errands, picks the nearest candidate for each, orders them into a loop and shows the verdict and route on the map.

The live demo uses a fixed public landmark as its start point, so it never needs or reveals my real location. You can also tap the map to set your own start.

I tested the app on my phone in demo mode, but I did not complete a timed walk to compare predicted and real minutes. That comparison is the next thing I'd do.

Type a messy errand list ("pharmacy, post office, buy rice, return parcel"). Get one loop you can walk or bike, with times for walking, cycling and driving, and a one-line verdict such as "3 stops, 17 min on foot (4 min by car). Skip the car."

A small open-weight model (Gemma 3 1B, run by llama.cpp) turns the text into structured errands OpenStreetMap data (via Overpass) finds the nearest matching places. OpenRouteService routes the loop It is a phone-first PWA built for the DEV Hacktoberfest Open-Source AI Challenge, Week 1 "Touch Grass". The model needs no API key and no per-call cost, and with the bundled model no AI vendor sees your text.

Built with an AI coding agent (OpenCode, model claude-sonnet-5-5) from a human-written product spec (`PRD.md`).

```
Phone (PWA, Leaflet)
  | POST /api/plan {text, lat, lng}
  v
FastAPI
  |- parser:  LLM (llama-server,
```

…
MIT licensed. The README covers running it locally (Ollama or llama.cpp), the Docker and Render deployment, the privacy data flow, the full eval tables and the known limitations. The project was started on October 11, inside the challenge window, and any commits after the deadline are listed in the README.

**The open pieces**

`llama-server` in production, Ollama on my laptop. The app talks to either through an OpenAI-compatible endpoint, so the model is a set of environment variables, not a dependency.
**One job for the model.** The LLM does exactly one thing: map each errand to one of ten fixed categories (pharmacy, post office, groceries, ATM, bakery, hardware, laundry, cafe, clinic, market) or `other`. The output is constrained to a JSON schema, so a 1B model can't wander. On the first attempt, 100% of its outputs were valid JSON in every run. Categories map to OSM tags, so the model never invents a place. It only picks a type of place, and the real places come from map data.

**Three layers around the model.** A small model is not trustworthy by itself, so I put deterministic code around it:

**How accurate is it?** I wrote a dev set of 26 cases (57 expected errands, English and Indonesian) and later added a held-out set of 15 cases (26 errands) that I wrote after freezing the prompt and keyword list. A test pins a hash of the frozen config, so I couldn't tune against the held-out set by accident. Intervals are 95% Wilson intervals, because these samples are small.

| Set | Parser | Accuracy | 
|---|---|---|
| Dev (57 errands) | raw model, Ollama | 82.5% [71-90] | 
| Dev | raw model, deployed container | 82.5% | 
| Held-out (26 errands) | raw model, Ollama | 96.2% [81-99] | 
| Held-out | model + keyword merge, container | 92.3% | 
| Held-out | + grounding check ‡, container | 96.2% | 

‡ I built the grounding check *after* seeing a held-out failure, so that row only shows it fixes the case that motivated it. It is not an independent result.

Some things I'd rather say than hide:

**Latency on Render** (1 CPU / 2 GB, demo point, three errands, Overpass cache pre-warmed at startup):

| Request | Total | Model call | 
|---|---|---|
| First after deploy | 6083 ms | 5348 ms | 
| Warm | 5896 ms | 5172 ms | 
| Warm | 5811 ms | 5061 ms | 

This is n=3 with one input. The first request is the first after a fresh deploy, not a cold boot from suspend, and I did not measure cold boot, a six-errand list, or a location the cache hasn't seen. Those will be slower, since an uncached location needs a live Overpass query.

**Things that went wrong along the way**

`start.sh` now sizes threads to the quota.
**Privacy, honestly.** Errand text never goes to an LLM vendor, because the model runs inside my own container. But location does leave your device: the server sends your start point rounded to three decimals to Overpass, and the exact start to OpenRouteService. The server doesn't store or log coordinates or errand text (I grepped the container logs to check), and the app has no analytics and no cookies.

**AI use.** I built this with Claude Code. That tool is not open source, and I don't count it as the open-source AI in this project. The open-source parts are the model, the runtime and the map data.

Here is what open weights gave me that a closed API wouldn't have:

The honest cost is accuracy. A larger hosted model would probably handle messy Indonesian better and would not have invented a pharmacy errand out of "main game." With a 1B model, you have to build the guardrails yourself: the schema, the keyword layer, the grounding check and the fallback. For a narrow task like this one, I think that trade is worth it. It's also the part of the build I learned the most from.
