This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
I spend most of my day on a screen. So I built Touch Grass, an app that gives me a reason to leave it.
Every morning, you get one outdoor task. "Find moving water." "Say hello to three people you pass." "Pick up ten pieces of litter." You can do it any time that day. Then you take a live photo as proof. A small AI on my laptop looks at the photo, and plain code decides if you passed.
The tasks are not only about nature. Some are about people: ask a shop owner how long they've been open, or ask someone for their favorite place to eat. The goal is to go out, look around, and talk to someone, not just count steps.
Today's task. Anytime today. No timer.
What it does:
It's for people like me who want a small push to go outside, and a little game to keep it going.
No signal? The photo is saved on the phone and waits.
Back online. The laptop checked the photo and the task is done.
Single-user app that gives you one daily outdoor task (nature, movement, exploration, social, creative, community). You submit a live photo as proof from a mobile app; a laptop server verifies it with a local VLM served by Ollama.
See AGENTS.md for the full spec, hard rules and game/verdict rules.
app/)
ollama pull qwen2.5vl:3b
eas build)
uv sync
copy .env.example .env # Windows; use `cp .env.example .env` on macOS/Linux
… The AI in the app is open and runs on my own laptop. It's an everyday machine: an RTX 3050 with 4 GB of video memory and 16 GB of RAM.
The AI. I used Qwen2.5-VL (3B) through Ollama. It looks at the photo and answers small yes/no questions, like "Is a leaf visible?" It also writes a short note about what it sees, and that note becomes the hint you read in the app.
The rule that made it work: the AI never decides pass or fail. Small AI models like to say yes. So plain code makes the call. You pass only if every required question gets a "yes". That also means nobody can sweet-talk the AI into a pass.
Choosing the model. I ran a quick test on 7 photos:
| Model | Right answers | "Not sure" answers | Speed per photo |
|---|---|---|---|
| Qwen2.5-VL 3B | 12 of 12 | 2 | 3 to 6 seconds |
| Gemma 3 4B | 8 of 10 | 4 | 4.5 to 6.5 seconds, but 71 seconds on the first photo, and it crashed once |
Gemma also gave "no" answers while its own note said the thing was visible. That's a tiny test, not a benchmark, but it was enough to pick Qwen. Swapping the model took one line.
The rest of the build:
Being honest about how I built it: I used coding agents (OpenCode and Cursor) to write most of the code, from detailed prompts and a rules file. Those run on hosted models. The AI that runs inside the app is the local one. I built the Android file with Expo's cloud build.
What doesn't work well yet:
The app counts its own screen time. That 7 minutes was a testing day.
You decide which kinds of tasks you want more or less of.
A tiny photo was stopped in 0.6 seconds. My test photo was only 30 KB. The photo check caught it before the AI even looked, and that's the point of checking first. The message said "move a little closer", which was wrong, because the problem was resolution, not distance. I changed it to say the resolution is too low.
Gemma disagreed with itself. One of the two models I tried said "no" to "is this outdoors?" while its own note said "sky is visible". It also took 71 seconds on the first photo and crashed once. Qwen got every question it answered right and was much faster. That's why I picked Qwen. It was a small test, but the difference was clear.
I typed the model name wrong, twice. My first try at a third model failed with a "not found" error. The tag had a typo. I fixed it and moved on, because Qwen was already doing the job.
A photo of a screen passed the "outdoors" question. I photographed a monitor showing a tree, and the AI said "outdoors". When I asked directly "is this a screen?", it got it right. So the model can tell, I just wasn't asking. For now, spotting screens is only logged. It's a clear next step.
The same note appeared for every question. The small model describes the photo once and pastes that note under each answer. So some hints are less useful than I'd like. A fix would be to ask one question at a time, which costs a few more seconds.
One extra retry made a check take 13 seconds instead of 5. My rule that every note must be at least four words sometimes made the model try again. It's a fair trade for better hints, but I watch it.
Uploads failed with "Unsupported FormDataPart". Expo's network code didn't accept the usual way of sending a photo. The server never saw a request. The app also said "couldn't reach your server" for every kind of error, which was misleading. I switched the upload to a different method and made the error messages say what really went wrong.
My history thumbnails were blank. The image requests didn't carry my token, so my own server said "401 unauthorized" to my own app. I now download each thumbnail with the token and keep a copy on the phone.
The app showed "done" when the server said it wasn't. When I deleted a test photo on the laptop, the phone still remembered it. Today's task showed as finished with no way to upload. I made the phone check with the server and drop items the server doesn't know about, and I added a button to clear the local cache.
The Close button did nothing after an "unsure" result. I tapped it three times. It only went away when I switched tabs. The cause was the screen being redrawn from fresh data, which brought the card back. Closing now takes effect immediately and stays closed.
Expo Go can't open the app in airplane mode. It loads the app from my laptop, so with no connection it can't start. The offline queue works once the app is open. For real use I needed a standalone app that doesn't depend on my laptop.
An app that tells you to go outside shouldn't send your photos to a company on the way. This one doesn't.