Can a Vision Model Tell If You Actually Went Outside? A developer built Touch Grass Quest, a daily, mobile-first web app that assigns one outdoor photo task per day and uses an open-weight vision model (Gemma, after switching from Gemini 2.5 Pro) to judge whether the submitted photo matches the task and was actually taken outdoors. The Go backend keeps the verdict logic in code rather than the model, requiring both a match and an outdoors flag, and the prompt explicitly rejects photos of screens or printed pictures to block the obvious cheat of photographing a monitor. The developer reports that in early tests the app rejected a photo of a berry bush displayed on a laptop screen, though they caution this is a design sanity check rather than a benchmark. This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass https://dev.to/challenges/hacktoberfest-week1-2026-10-05 Touch Grass Quest is a daily outdoor photo task that is designed to be over in seconds. You open the page and read one small task, like "Find a berry on a bush." Then you put the laptop down and go outside. You take one photo, and a vision model tells you in one line whether it matches. After a pass, the app says "Done for today. Go touch more grass." and stops. There is no feed, no notifications, and no leaderboard. The most the app does after a pass is show a streak counter. I wanted the screen to be the shortest part of the experience, and the best way to do that was to give the app nothing else to offer. It's for people like me who spend most of the day at a laptop and keep meaning to go outside. I haven't deployed a hosted version. The app runs locally, and the repo includes a Dockerfile for anyone who wants to host it. The screenshots here are from my own testing. Touch Grass Quest is a daily, mobile-first web app built for the DEV Challenge. It encourages users to put down their screens and interact with the real world by giving them one unique outdoor photography task every day. Upload your photo, and an open-weight Vision AI will determine if you successfully found the item outdoors. Build your streak, and go touch some grass net/http , plus a small front end in plain HTML, CSS, and JavaScript. There's no framework and no database. tasks.json . The task of the day is picked deterministically from the date, so everyone gets the same one. Each task carries a safety hint. The berry task says "look for wild berries, but don't eat them." VISION MODEL setting exists because a text-only model can silently ignore an image, so image checks can be routed to a model that actually sees them. I wrote a small test tool cmd/visiontest to confirm that Gemma really reads the photo and isn't answering from the task text alone. The first design decision was to keep the verdict out of the model. The model returns one strict JSON object, with a match flag, an outdoors flag, and a short reason . Go then makes the call: a photo passes only if it matches the task and looks outdoors. If it matches but looks indoors, the reply is gentle "Looks like it's indoors or a screen, take it outside" . The prompt asks the model to be fair but honest, to accept reasonable interpretations, to never invent objects, and to treat a photo of a screen or a printed picture as not outdoors. I wrote that last rule after thinking about the most obvious way to cheat, which is photographing a nice forest on your monitor. The failure paths are handled in code too: