Grass Gate: my feeds only unlock after an open-weight model sees me outside A developer built Grass Gate, an open-source doomscroll lock that blocks distracting sites via a Chrome Manifest V3 extension and only unlocks them after the user walks to a real nearby location and photographs it. The phone PWA uses OpenStreetMap's Overpass API to pick quests and runs the open-weight CLIP ViT-B/32 vision model on-device through transformers.js (about 155 MB, cached) to zero-shot verify the photo is outdoors and matches the target feature, then issues a one-time 6-digit code that unlocks feeds for 30 minutes. The app is plain HTML, CSS and JavaScript with no bundler or backend, and the code is available on GitHub. This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass https://dev.to/challenges/hacktoberfest-week1-2026-10-05 Grass Gate is a doomscroll lock that only opens outdoors. A browser extension blocks the sites I lose hours to. When I hit one, I don't get a guilt-trip timer. I get a quest. I open Grass Gate on my phone, and it picks a real place near me from OpenStreetMap: 💧 Mallard Lake is 330 m south. Walk there and photograph the water. I pocket the phone and walk. When I arrive, it buzzes and chimes. I take the photo, and an open-weight vision model running on the phone checks two things: am I actually outdoors, and is that actually water? If both pass, I get a one-time 6-digit code. I type it into the blocked page, and my feeds open for 30 minutes. Then they lock again. The screen is only the lock. Everything that counts happens outside. | Today | Quest complete | Field journal | |---|---|---| It's for anyone whose thumb opens a feed before their brain has decided to, which is me. A few things make it more than a gimmick: Try it on your phone: grassgate.vercel.app https://grassgate.vercel.app/ Open it in your phone's browser, allow location and camera, and tap Find me a quest . Add it to your home screen to install it. After the first photo check, the vision model is cached and works without signal. The app works on its own: quests, the on-device photo check, the journal, streaks and stamps. The browser extension adds the lock, and the 6-digit code from a finished quest unlocks it. Want to try it at your desk? Turn on Demo mode in Settings to skip the walking distance check. Your feeds stay locked until you go outside. ▶ Try it: grassgate.vercel.app https://grassgate.vercel.app/ open it on your phone A browser extension blocks your doomscroll sites. To get back in, open Grass Gate on your phone. It picks a quest from real places near you "Elk Glen Lake is 430 m southeast, go photograph the water" , then walks you there. An open-weight vision model on your phone checks the photo, and you get a one-time 6-digit code that unlocks your sites for 30 minutes. The screen is only the lock. The whole experience happens outside. On your phone PWA Repo: github.com/Rahul-Roy-Hub/grassgate https://github.com/Rahul-Roy-Hub/grassgate The whole app is plain HTML, CSS and JavaScript: a PWA you can install, about 100 KB of app code, with no bundler and no backend. The extension is a Chrome Manifest V3 extension. Laptop Chrome extension Phone PWA declarativeNetRequest blocks feeds GPS → OpenStreetMap Overpass → nearby features "Go touch grass" page → quest: tree / water / park / bench / art / view optional local LLM writes the quest text shared pairing secret walk; camera unlocks within ~40–150 m ◀──── typed once, never sent anywhere ────▶ CLIP ViT-B/32 checks the photo on-device checks the 6-digit code offline ◀── code ── pass → one-time code valid 5–10 min 1. CLIP ViT-B/32 via transformers.js: the referee. The open-weight CLIP https://huggingface.co/Xenova/clip-vit-base-patch32 model runs in the browser through transformers.js https://github.com/huggingface/transformers.js about 155 MB, downloaded once and cached . Zero-shot classification means I didn't train anything: the classifier's "knowledge" is just a list of English labels, the same trick described in Is this even a valid card? Zero-shot image classification model in a lambda container https://dev.to/aws-builders/is-this-even-a-valid-card-zero-shot-image-classification-model-in-a-lambda-container-58oj . Mine runs on the phone instead of in a Lambda. Each photo gets two checks: My first version added up the scores of every synonym: water + a river + a pond + a lake + a fountain . During end-to-end testing with a simulated camera, a photo of a beach towel and a book on sand passed the water quest at 41%. Nothing in the photo was water. With five water labels among thirteen, the target collects about 38% of the probability before the model has looked at anything. My cutoff was 35%. The fix uses the fact that softmax keeps logit differences, so ln p a / p b = logit a − logit b : js // Best target label must rank top-3 AND beat the median decoy by ~2.7×. const logit = Object.fromEntries target.map r = r.label, Math.log r.score ; const best = Math.max ...quest.labels.map l = logit l ; const decoyLogits = decoys.map l = logit l .sort a, b = a - b ; const margin = best - decoyLogits Math.floor decoyLogits.length / 2 ; const rank = target.findIndex r = quest.labels.includes r.label + 1; const targetOk = margin = 1.0 && rank <= 3; With 18 decoys "sand", "a pavement", "an animal", "a book"… , all 7 of my sample photos now land correctly. The towel fails with "Couldn't spot water. It looked more like a book." A butterfly on a flower and a dog on grass pass. Seven photos is a small set, so these thresholds will keep moving as I test outside. Because they're plain numbers in my own code, I can move them. The app shows its working, too. Under every result, "What the on-device model saw" lists the real scores. Here, "an animal" is CLIP's top guess, but grass still ranks 2 and clears the margin easily: 2. OpenStreetMap via the Overpass API: the quest-giver. A single Overpass query pulls trees, water, parks, woods, benches, playgrounds, viewpoints and public art within your chosen radius. Quests are at least 120 m away so you actually walk, and the last map is cached so it keeps working without signal. If nothing is mapped nearby, or you're offline, you get "anywhere" quests: sky, leaf, flower, grass. 3. An optional local LLM via WebLLM: the narrator. With WebGPU, an open-weight model running in the browser Gemma 2 2B, Qwen2.5 0.5B or Llama 3.2 1B, swappable in settings rewrites each quest so it reads like a friend nudging you outside. The LLM never decides anything that matters. Quest types, CLIP labels and geofences are deterministic. The model only gets facts I already have place name, distance, direction, season and writes the words. A tiny or swapped model can make the text worse, but it can't send you to a place that doesn't exist. 4. No server at all. The phone and the extension share a secret once, by link or by typing a 16-character code. Unlock codes are HMAC-SHA256 of that secret and a 5-minute time window, the same idea as TOTP, checked offline on both sides. Each quest gives exactly one code: the extension refuses reuse and rate-limits wrong guesses. Because the core of this app is a daily, geotagged photo of where I live and walk. Grass Gate isn't something I'd ship on a closed vision API. | | Grass Gate open, on-device | Same idea on a closed API | |---|---|---| | Your photos | Never leave the phone | A daily photo trail of your neighbourhood on someone else's server | | Signal | Works on a trail after one model download | Every check needs a connection | | Cost | $0 per check, no keys, no rate limits | Paid per image, plus a backend to hide the key | | Server | None | Needed for state and for the API key | | Control | I read the model's scores, retuned the thresholds, and fixed the towel bug in an afternoon | A yes/no from a black box I can't calibrate | The last row is where open won the hardest. The beach-towel bug was only fixable because I could see the full probability distribution and work back to the logits. A closed "does this photo contain water?" endpoint would have said yes , and I'd have had no way to know why. There are honest trade-offs. The first CLIP download is about 155 MB, and ViT-B/32 is not the sharpest model available. But it's small enough to run on a phone, and good enough to tell a lake from a towel.