I built a bingo card you can't fill from your desk A developer built Proof of Grass, a daily outdoor bingo card web app that uses CLIP ViT-B/32 quantized to 8-bit and run in-browser via transformers.js to verify photos of outdoor items. Testing against 131 Wikimedia Commons photos showed a raw-similarity threshold approach filled only 56 of 102 correct squares, while softmax over 25 outdoor labels reached 96 of 102 but wrongly accepted all 29 indoor photos, including a monitor classified as grass. This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass https://dev.to/challenges/hacktoberfest-week1-2026-10-05 "Touch grass" is usually an insult. I wanted to turn it into a game you can only win outside. Proof of Grass is a daily bingo card for the outdoors. Every morning there is a new 3×3 card: a fallen leaf, moss, a snail, a park bench, a puddle with something reflected in it. The centre square is always grass. Tap a square, take a photo, and a vision model running on your phone decides whether you actually found the thing. Complete a line, get a bingo. Share your grid like a Wordle score. Everyone gets the same card on the same day, so "did you find the snail?" is something two people can actually say to each other. There is one free swap a day for the square your neighbourhood cannot provide, and a streak that only counts days you filled the grass square. It is for anyone whose screen-time report has started to feel like a jump scare: a ten-minute walk with a reason to look closely at things, for people who need a reason. It also works as a Saturday morning activity with kids, because "find a spider web" is a very good reason to look properly at a hedge. 🔗 vanshajpoonia.github.io/proof-of-grass https://vanshajpoonia.github.io/proof-of-grass/ No sign-up, no account. It is built for a phone, so open it on one if you can. If you are judging this from a desk no shame, so was I , every square has two sample photos: the real thing , and a laptop on a sunny balcony . The first fills the square. The second gets told: "That is a keyboard. Spiritually adjacent, physically not." The first photo downloads the model, 89 MB, once. After that, put your phone in airplane mode. It keeps working. A daily outdoor bingo card that only a real photo can fill. Every morning there is a new 3×3 card: a fallen leaf, moss, a snail, a park bench, and grass in the middle, always. Tap a square, take a photo, and a vision model running on your phone decides whether you actually found it. Fill a line for a bingo. Share your grid. Same card for everyone, every day. Live: https://vanshajpoonia.github.io/proof-of-grass/ https://vanshajpoonia.github.io/proof-of-grass/ Static files, no build step, no backend. MIT. The judge is CLIP ViT-B/32 https://huggingface.co/Xenova/clip-vit-base-patch32 , OpenAI's open-weight image/text model, quantised to 8-bit and running in the browser tab through transformers.js https://github.com/huggingface/transformers.js on ONNX Runtime Web. Everything else is plain HTML and JavaScript and a service worker. CLIP compares pictures with sentences, so "is this a pinecone?" sounds like one line of code. It was not, and the story of why is the most useful thing I learned this week. To keep myself honest I collected 131 photos from Wikimedia Commons : 102 of the 25 outdoor items, and 29 that should never fill a square laptops, keyboards, bedrooms, lunch, car dashboards . The test suite runs the real model over all of them, so every attempt below has a number. Attempt 1: "is it close enough?" Compare the photo to the text a photo of a pinecone , accept above a threshold. I swept for the best threshold and it still only filled the right square for 56 of 102 photos. CLIP's raw similarity scores sit in a narrow band around 0.2 to 0.3, and that band means different things for different words. Attempt 2: "is it more pinecone than anything else on the card?" Softmax across all 25 outdoor items, accept if the claimed one wins. 96 of 102. I was delighted, briefly, until I ran it on the indoor photos. All 29 of them filled a square. A softmax has to spend 100% of its probability somewhere. If the only options are outdoor things, a photo of a keyboard is forced to be something outdoors, and it was. A computer monitor showing a test pattern was grass . A laptop next to a ruler was an insect , with 58% confidence. Two different farmhouse bedrooms were park benches . A backlit keyboard was a dog . Attempt 3: give it somewhere else to look. I added nine classes that are never squares: a screen, a keyboard, a room, food, a face, the inside of a car, a houseplant, a wall, a page of text. Now a photo of a laptop has a better answer than "probably moss". Indoor photos that filled a square: 0 of 29 . Accuracy on outdoor photos did not move. Attempt 4: let a photo be two things. A bee on a flower should be able to fill insect or flower . Accepting the claimed item if it is in the top two with at least 15% of the probability took it to 99 of 102 , still 0 of 29 indoors. | Approach | Right square | Indoor photos that ticked a square | |---|---|---| | Cosine ≥ best threshold | 56 / 102 | 0 / 29 | | Softmax over outdoor items | 96 / 102 | 29 / 29 | | + "elsewhere" classes | 96 / 102 | 0 / 29 | | + prompt ensembling | 96 / 102 | 0 / 29 | | + top-2, p ≥ 0.15 shipped | 99 / 102 | 0 / 29 | Two honest notes on that table. Prompt ensembling , averaging several phrasings like a close-up photo of moss , is the standard CLIP trick and it did nothing measurable here. I kept it only because it is free at runtime. And the three misses are two small feathers lying on the ground and a puddle in a farm track: things that take up a few percent of the frame. CLIP looks at the whole picture, so a tiny feather on a big lawn is, reasonably, a photo of a lawn. That is why a "no" in the app says what it saw instead and tells you to get closer, rather than just failing. A few other decisions that came out of measuring rather than guessing: Because of what is in the photos. A photo taken outside on your phone is a small map of your life: your street, your park, the route you walk every evening, your kid in the corner of the frame. Sending that to a hosted vision API so it can confirm that, yes, this is moss, is a ridiculous trade. With open weights running in the tab, "your photos never leave your phone" is not a line in a privacy policy that can be revised. It is a property of where the code runs, and you can check it yourself: open the network tab, take a picture, and watch nothing happen. Three more things a closed API would not have given me: I built this with Claude Code as a pair programmer. I didn't save the session to DevRelay, so here is the next best thing: everything in this post can be re-run. git clone https://github.com/VanshajPoonia/proof-of-grass cd proof-of-grass && npm install && npm test That downloads the 131 test photos credits in tests/fixtures.json , runs all five approaches through the real model, and prints the table above. MODEL=Xenova/clip-vit-base-patch16 npm test reproduces the bigger-model comparison.