{"slug": "can-a-vision-model-tell-if-you-actually-went-outside", "title": "Can a Vision Model Tell If You Actually Went Outside?", "summary": "A developer built Touch Grass Quest, a daily, mobile-first web app that assigns one outdoor photo task per day and uses an open-weight vision model (Gemma, after switching from Gemini 2.5 Pro) to judge whether the submitted photo matches the task and was actually taken outdoors. The Go backend keeps the verdict logic in code rather than the model, requiring both a match and an outdoors flag, and the prompt explicitly rejects photos of screens or printed pictures to block the obvious cheat of photographing a monitor. The developer reports that in early tests the app rejected a photo of a berry bush displayed on a laptop screen, though they caution this is a design sanity check rather than a benchmark.", "body_md": "*This is a submission for the [Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass](https://dev.to/challenges/hacktoberfest-week1-2026-10-05)*\n\n**Touch Grass Quest** is a daily outdoor photo task that is designed to be over in seconds. You open the page and read one small task, like \"Find a berry on a bush.\" Then you put the laptop down and go outside. You take one photo, and a vision model tells you in one line whether it matches. After a pass, the app says \"Done for today. Go touch more grass.\" and stops.\n\nThere is no feed, no notifications, and no leaderboard. The most the app does after a pass is show a streak counter. I wanted the screen to be the shortest part of the experience, and the best way to do that was to give the app nothing else to offer.\n\nIt's for people like me who spend most of the day at a laptop and keep meaning to go outside.\n\nI haven't deployed a hosted version. The app runs locally, and the repo includes a Dockerfile for anyone who wants to host it. The screenshots here are from my own testing.\n\n**Touch Grass Quest** is a daily, mobile-first web app built for the DEV Challenge. It encourages users to put down their screens and interact with the real world by giving them one unique outdoor photography task every day.\n\nUpload your photo, and an open-weight Vision AI will determine if you successfully found the item outdoors. Build your streak, and go touch some grass!\n\n`net/http`), plus a small front end in plain HTML, CSS, and JavaScript. There's no framework and no database.`tasks.json`. The task of the day is picked deterministically from the date, so everyone gets the same one. Each task carries a safety hint. The berry task says \"look for wild berries, but don't eat them.\"`VISION_MODEL` setting exists because a text-only model can silently ignore an image, so image checks can be routed to a model that actually sees them. I wrote a small test tool (`cmd/visiontest`) to confirm that Gemma really reads the photo and isn't answering from the task text alone.\nThe first design decision was to keep the verdict out of the model. The model returns one strict JSON object, with a `match` flag, an `outdoors` flag, and a short `reason`. Go then makes the call: a photo passes only if it matches the task **and** looks outdoors. If it matches but looks indoors, the reply is gentle (\"Looks like it's indoors or a screen, take it outside\").\n\nThe prompt asks the model to be fair but honest, to accept reasonable interpretations, to never invent objects, and to treat a photo of a screen or a printed picture as not outdoors. I wrote that last rule after thinking about the most obvious way to cheat, which is photographing a nice forest on your monitor.\n\nThe failure paths are handled in code too:\n\n`<thought>` tags ahead of the real answer. I strip those before parsing, and I treat an unclosed tag as a failure instead of showing half a thought process.\nThe obvious way to cheat is to photograph a nice picture on your monitor, so I tried exactly that. I pointed the camera at a picture of a berry bush on my laptop screen and submitted it.\n\nIn my first test runs, which used Gemini 2.5 Pro before I switched to Gemma, the app rejected photos like this. The model recognized the berries and still refused to pass them, because it noticed laptop bezels in the frame. That's the behavior I wanted. A couple of rejected photos isn't an accuracy number, though, and I changed models afterward, so I treat it as a sanity check on the design and not a benchmark.\n\n`'` where an apostrophe should be. My code was building the message badly and escaping it wrongly.\nI'll be careful here, because I didn't compare against a closed model on the same photos, so I can't claim open was more accurate. What open gave me:\n\n`LLM_BASE_URL`, `LLM_MODEL`, `.env` settings.\nA proper field test on real walks, with more people than just me and a measured hit rate. A larger photo set in the evaluation harness, including the failing cases above. And a hosted deployment so it doesn't depend on my laptop.\n\nI built this with Antigravity, working from a detailed spec that asked for tests, strict parsing, and an evaluation harness. I then ran it, tested it by hand, and found the bugs above myself.", "url": "https://wpnews.pro/news/can-a-vision-model-tell-if-you-actually-went-outside", "canonical_source": "https://dev.to/ghanshyam_jha/can-a-vision-model-tell-if-you-actually-went-outside-1lne", "published_at": "2026-10-10 08:48:49+00:00", "updated_at": "2026-10-10 09:10:25.285039+00:00", "lang": "en", "topics": ["computer-vision", "ai-tools", "generative-ai", "developer-tools"], "entities": ["Touch Grass Quest", "Gemma", "Gemini 2.5 Pro", "DEV Challenge", "Hacktoberfest"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/can-a-vision-model-tell-if-you-actually-went-outside", "markdown": "https://wpnews.pro/news/can-a-vision-model-tell-if-you-actually-went-outside.md", "text": "https://wpnews.pro/news/can-a-vision-model-tell-if-you-actually-went-outside.txt", "jsonld": "https://wpnews.pro/news/can-a-vision-model-tell-if-you-actually-went-outside.jsonld"}}