# I walked my neighborhood with the phone in my pocket, and a local Gemma wrote the report

> Source: <https://dev.to/darrkkens/i-walked-my-neighborhood-with-the-phone-in-my-pocket-and-a-local-gemma-wrote-the-report-7f2>
> Published: 2026-10-06 19:44:28+00:00

*This is a submission for the [Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass](https://dev.to/challenges/hacktoberfest-week1-2026-10-05)*

**Bairro em Ação** ("Neighborhood in Action") is a walk-first app for reporting what's broken in your neighborhood: potholes, cracked sidewalks, piles of rubbish, a bent traffic sign, a broken park bench.

I live in Joaçaba, a small city in Santa Catarina, Brazil. Everyone notices these things on the way to the bakery, but few people report them, because reporting means stopping, typing, picking a category, finding the address, and writing something the city hall will take seriously. On a sidewalk, with one hand, that's too much.

So the app does as little as possible while you walk:

Once the photos reach your computer, an open-weight model, **Gemma 3 4B running locally through Ollama**, looks at each one in the background. It suggests a category, a short title and a description. When you get home, you go through the suggestions, confirm or fix each one, and share a report. It's a single HTML file with the photos, a map of your route and every problem grouped by category, ready to send to the city's ombudsman channel or a residents' association group chat.

It's for people who already walk their streets: neighbors, residents' associations, a parent on the school run. The goal is that the screen is the shortest part of the walk.

There's no hosted demo on purpose: the whole point is that the model runs on your own computer (more on that below). Here is the flow on a phone. These screenshots use freely licensed test photos from Wikimedia Commons. The photos from my real walk stay on my machine, except three shots of one sidewalk that you'll see further down.

**Walking.** One button, and a running list of points with their AI status ("Na fila da IA" = queued, "IA analisando" = analyzing):

**A new point.** Photo, an optional note, quick chips, location captured while you type:

**Review at home.** Gemma's suggestion comes prefilled. You confirm it or change it:

**The route.** Every photographed point is numbered on an OpenStreetMap map of the walk:

**The report.** One self-contained HTML file, no scripts, opens offline, prints to PDF:

**Walk your neighborhood, photograph what needs fixing, and get an organized report from an open-weight model running on your own computer.** · *Caminhe pelo bairro, registre problemas nos espaços públicos e transforme fotos em um relatório com IA aberta.*

Bairro em Ação is a mobile web app (interface in Brazilian Portuguese) with a Go API, PostgreSQL and **Gemma 3 4B running locally through Ollama**. Its core is the **walk mode**: during the walk you only take a photo, optionally type a word and keep going. The AI reads each photo in the background. When the walk ends, you review the suggestions and share a report with photos, places and descriptions.

`net/http`), PostgreSQL via `pgx`.` gofmt`, `go vet`, `go test -race` against a real PostgreSQL, and the frontend build. Contributing guide, issue templates (including one for "the AI suggestion was wrong") and a security policy.

```
Phone (React PWA)                          Your computer
camera · note · GPS route        PUT       Go API ─► PostgreSQL (walks, points, route)
resize + strip EXIF          ─────────►      │
IndexedDB outbox (offline)                   ├─► CLIP (ONNX, ~0.3 s): off-topic? same spot?
service worker                               │
review · Leaflet map · share  ◄─────────     └─► worker ─► Gemma 3 4B via Ollama (~80 s)
                                polling                     JSON schema, validated in Go
```

A walk is created on the phone with a client-side UUID, so it starts even with no signal. Each point goes into an IndexedDB outbox and uploads by itself when the server is reachable. Every upload is idempotent, so a retry never creates a duplicate. A service worker keeps the app opening offline. Before leaving the phone, each photo is re-encoded to a 1600 px JPEG, which also strips the EXIF metadata, including the camera's own GPS tag.

The server is my notebook. One script (`servidor-celular.sh`) starts PostgreSQL and Ollama, builds the app and serves everything over HTTPS on the home Wi-Fi, with a small local certificate authority so the phone allows GPS and offline storage. Open the app once at home, walk with no connection, and the queue drains when you're back.

A single background worker takes one point at a time (`FOR UPDATE SKIP LOCKED` in PostgreSQL) and sends the photo plus the note to Gemma through Ollama's `/api/chat`. Three things make a 4B model reliable enough for this:

`format` field takes a JSON schema. The category is an `enum` of the five categories and the confidence is `alta`/` media`/` baixa`, so the model can't invent a sixth category. Go validates the answer again (enum, lengths) before storing it.`question` field, and the prompt asks Gemma to fill it with one short question when its confidence is low. The review card shows that question. Answer it and ask for a new analysis, or fill the point in by hand. If the model is unsure but forgets to ask, Go adds a default question itself.
The AI's answer (`ai_*` columns) and the person's confirmed values are stored separately, along with the model name. A late answer for a point the person already re-queued is discarded (`WHERE ai_status = 'running'`). Date and location always come from the app, never from the model.

On October 5 at 17:13 I went out for a walk in Flor da Serra, a neighborhood of Joaçaba, with the notebook at home as the server:

| Walk | 17:13 to 17:20, about 480 m of GPS route (24 positions) | 
| Points recorded | 9, all with location: 3 potholes, 4 broken sidewalks, a damaged traffic sign, litter | 
| AI category kept after review | **9 of 9** , all suggested with high confidence | 
| Analysis | The nine uploads reached the notebook together at 17:20, as the walk ended. Gemma worked through them one at a time until 17:33: 66–92 s per photo on a GTX 1050 (3 GB) | 

Nine photos aren't a benchmark, but the whole loop worked outdoors: capture, analyze, review, report with map. During the walk I never waited for the AI, which was the whole design bet.

Then I opened the report. This is part of it, exactly as the app generated it that afternoon. I only hid the coordinates:

Points 3, 4 and 5 are the same broken sidewalk. I photographed it three times in 28 seconds, at 17:16:04, 17:16:17 and 17:16:32. The phone placed the three photos 11 and 20 m apart, with a GPS error of ±23 to ±31 m. They were closer to each other than the GPS could tell apart.

The app treated them as three problems. Gemma analyzed each one from scratch: 81, 92 and 83 seconds, more than four minutes of a small GPU on one sidewalk. And because each analysis started from nothing, it wrote three different descriptions: a short one, one about elderly people and reduced mobility, and one about an inspection cover. Sent as it was, the report would give the city hall three complaints about one sidewalk, each worded differently.

Looking at that report, my idea was simple: some kind of image identification should run *before* the recognition, as a pre-filter for the AI. Photos of the same spot, taken seconds and meters apart, should be grouped, and duplicates dropped, before Gemma ever sees them.

Testing that day gave me one more reason for a pre-filter. A profile-picture placeholder sent with the note "Lixo acumulado" came back as *rubbish, high confidence*: Gemma trusted the note over the photo, despite the rule in the prompt. A photo that shows no problem at all shouldn't reach Gemma either.

(One more thing the walk taught me, which no model can fix: the route stops recording when the screen is off, because browsers pause location updates for background pages. The walk screen says when GPS is paused, and where the track has gaps, the map connects the photographed points instead.)

The obvious first try didn't work. Classic image descriptors (color and texture histograms, edge directions) can't tell one grey street surface from another: the three sidewalk shots scored 0.87–0.93 against each other, but unrelated problems from the same walk also scored up to 0.93. I also tried asking Gemma to compare five photos in one request. It took 358 seconds and grouped none of them.

What works is a second, much smaller open model in front of Gemma: **CLIP ViT-B/32**, run inside the Go backend with ONNX Runtime, at about 0.3 seconds per photo on the notebook's CPU. Every upload now goes through it first:

Here is how the grouping rule sees the photos from that walk:

| Pair | Time apart | Distance (GPS error) | CLIP | Result | 
|---|---|---|---|---|
| 3 → 4, same sidewalk | 13 s | 11 m (±23 m) | 0.91 | grouped ✓ | 
| 3 → 5, same sidewalk | 28 s | 20 m (±23 m) | 0.89 | grouped ✓ | 
| 5 → 6, another broken sidewalk further on | 36 s | 51 m (±31 m) | 0.75 | kept apart ✓ | 
| 1 → 2, cracked asphalt vs. a patched drain | 21 s | 9 m (±17 m) | 0.76 | kept apart ✓ | 
| 1 → 3, a pothole vs. the sidewalk | 45 s | 35 m (±18 m) | 0.86 | kept apart ✓ | 

The last row is why CLIP alone isn't enough: to CLIP, a pothole and a sidewalk 35 m apart look about as alike (0.86) as two shots of the same sidewalk. Time and place have to agree too.

Applied to the walk, the rule folds points 4 and 5 into point 3 and keeps everything else apart. That's 7 Gemma analyses instead of 9, about three minutes less of GPU time, and the sidewalk becomes one entry with three photos:

The off-topic filter kept all nine photos from the walk, plus 8/8 problem photos from Wikimedia Commons and 90/90 reference photos of rubbish, discarded furniture and damaged signs. It let through 2 of 17 off-topic photos, both street scenes with nothing wrong in them. That's deliberate: a discarded photo is deleted, so doubtful cases go on to Gemma, which can ask the person.

A test walk on the phone with six uploads (the avatar, three shots of the sidewalk, the asphalt, a cat) now produces **2 points and 2 discards: 2 Gemma analyses instead of 6**. Grouping is never final. Extra photos can be split apart in review, and points that still look alike after the AI get a "same point as #3? 11 m and 13 s apart" suggestion with **Agrupar** / **São diferentes** buttons. The AI suggests, a person decides.

The neighborhood list for a city comes from OpenStreetMap through the Overpass API (looked up by the IBGE municipality code and cached in PostgreSQL, because public Overpass servers are often overloaded). The phone's location fills in the neighborhood through Nominatim, with coordinates rounded to about 100 m first: enough for the neighborhood, not the house. City search uses BrasilAPI and IBGE, and typing a neighborhood with no city picked searches Brazil-wide with Photon. The route map uses Leaflet and OpenStreetMap tiles. For the report, the Go backend draws a static map from cached tiles plus an SVG overlay, so the HTML has no scripts and opens offline.

**Street photos are other people's lives.** A photo of a pothole also shows houses, cars, license plates and whoever was walking by, and every point carries coordinates. With Gemma and CLIP running on my notebook, none of that goes to a server I don't control. The only things that leave the machine are open-data lookups: rounded coordinates to Nominatim, a city name to BrasilAPI, map tile requests. None of them receive a photo.

**It costs nothing to run at neighborhood scale.** A residents' association can analyze hundreds of photos a month with no per-image fee, no quota, no account and no API key to leak. For a volunteer group, that decides whether the tool gets used at all.

**No signal on the walk is fine.** Capture needs no network, and the AI runs when you're back on Wi-Fi. A cloud API wouldn't make the walk any better. It would just cost mobile data.

**I could change the parts that were wrong.** When the report showed three entries for one sidewalk, I didn't have to wait for a vendor to add deduplication. I put a second open model in front of the first, measured the thresholds on my own walk's photos and had it running that same night. The prompt and the categories are plain Go code that another city can edit. `OLLAMA_MODEL=gemma3:12b` swaps the model with no code change, and since every suggestion stores which model made it next to what the reviewer confirmed, a group can compare models on their own streets instead of trusting someone else's benchmark.

**Where open beat closed for me:** the local CLIP check takes 0.3 seconds and saves an 80-second Gemma call for every repeated or off-topic photo. On this walk alone, that's two of nine analyses. With a hosted vision API, I'd be paying for, and uploading, every one of those repeats.

**The honest trade-off:** a 4B model on a 3 GB GPU is slow (66–92 s per photo, 13 minutes for this walk), and it can still be fooled by a confident note, which is why CLIP filters first and a person confirms every point before it goes in the report. For this job, slow and private beats fast and uploaded: the walk never waits for the model, and a 13-minute queue after a 7-minute walk is fine.

I built Bairro em Ação with the help of an AI coding assistant (Claude Code), which I directed, reviewed and tested throughout. This post was drafted with AI assistance and reviewed by me before publishing. The walk, the validation numbers, the report excerpts and the screenshots come from running the real application. The test photos are credited in the repository README.
