Can You Find It? An open model hides one real thing in a place you know A developer built "Can You Find It?", an open-source game in which a local Gemma 4 E4B vision model running via Ollama on a laptop secretly picks one real object in a photo of a familiar place and issues a single-line clue, then verifies the player's close-up photo with FOUND IT, ALMOST or NOT QUITE. Before building the UI, the developer tested the core trick on 20 photos of public places, hand-labelling 184 candidate targets and finding that 88% of the model's bounding boxes contained the described object with 0% invented, and that roughly 1 in 2 photos yields a good round. This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass https://dev.to/challenges/hacktoberfest-week1-2026-10-05 I go to the same café every week. Same counter, same stool, same order. Last night a 4-billion-parameter model running on my laptop looked at a photo of that counter, picked one thing in it, and gave me one line: I found something hiding in plain sight. It took me two hints to find it: a small framed award on the tiled wall, the kind of thing you stop seeing the second time you walk in. My field note, typed while I was still a bit annoyed with myself: "A café I go to every week. Finally I've sat at that counter dozens of times and never noticed it " That's the whole game. Can You Find It? is I Spy, but the AI does the spying, and it only uses what is really in front of you. Every "go outside" app I could think of has the same problem: the app is the thing you look at. Here the screen does two small jobs, the clue and the check, and the place does the rest. The model can see the street. It can't walk down it. That part is yours. It's for anyone who walks the same streets every day and has stopped seeing them, and for anyone who'd like a reason to look up from their phone that doesn't come from their phone. A real round at my café, recorded on my phone. The model's wait, the hint wait and my typing are sped up; nothing else is edited. I played this one from a photo at home, so the "search" happens in the photo. In the street, that part happens with your eyes and the phone in your pocket. There's no hosted demo on purpose: the model runs on your own computer, and that's the point more on that below . Not every round is that good. In another photo, an old one of mine from a square in France, it picked an empty bike dock, one of five identical ones. Technically there, not much of a hunt. More on that below. I Spy, but the AI hides something real in a place you know — and checks that you found it. Take a photo of where you are, or pick one of a place you pass every day: your street, the way to the bakery, the view from a window. A local, open model Gemma 4 looks at it, secretly chooses one real thing in it, and gives you a single line: I FOUND SOMETHING. Someone wanted this place to remember something. Find it with your own eyes, now or next time you're there. You look at the place, not the screen. When you spot it, you snap a close-up — right away, or a five-second photo on your way past, sent whenever you like The model compares it with what it saw: FOUND IT, ALMOST or NOT QUITE. At the end you see both side by… Next.js 16, React 19, TypeScript, Node, Ollama https://ollama.com and Gemma 4 E4B. npm run doctor checks Ollama, the model and your memory, and prints a QR code for your phone. Everything runs on Gemma 4 E4B open weights, Apache 2.0 through Ollama , on my own laptop: local inference, no cloud API anywhere. The phone is just a browser on the same Wi-Fi; a small Next.js server on the laptop holds the rounds and talks to the model. The whole design is shaped around what a small open vision model can and can't do, so I started by measuring that. Before building any UI I wanted to know if the core trick works: can a model look at a wide photo of a real place and point at one specific, real, findable thing, without making it up? I took 20 photos of public places parks, squares, streets, gardens, a playground, a forest trail , asked the model for targets with boxes, cut every box out as a real crop , and labelled all 184 candidates by hand. No trusting the model's opinion of itself. What I learned: box 2d on a 0–1000 grid. With the final prompt, 88% of boxes held the thing it described, and 0% were invented. In the end, about 1 in 2 photos gives a genuinely good round, 1 in 5 gets an honest "I couldn't find anything I'd trust here" a forest trail has no plaques , and the rest are playable but mundane. photo ──► Gemma proposes 2 short targets with boxes ──► cheap filters: no areas, people, animals, vehicles, huge boxes, look-alikes ──► multiple-choice check on the real crop + "is a person at it?" ──► human-written clue line, chosen by rules ← shown right away ──► hints written in the background while you start looking ──► your close-up ──► FOUND IT / ALMOST / NOT QUITE The first version took a median of 47 seconds to show a clue. Two targets instead of three, shorter labels, a smaller crop for verification, and showing the clue before the hints are written brought it to 26 seconds. A tiny warm-up pass when you open the camera wakes the model while you frame the photo. Then I found the real problem: memory. The model needs about 9.5 GB, my laptop has 16, and with a browser, an editor and a few chat apps open, macOS starts swapping. The same call that takes 11–22 seconds took 27–45 seconds, and once sixteen minutes . npm run doctor now warns you when the computer is swapping, which is the most useful line of code in the project. I also tried sending a smaller photo 1536 px instead of 1920 . It's 25% fewer image tokens, but it lost exactly the best small targets, like a memorial plaque, so I kept 1920. To test the losing screen, I photographed my laptop screen instead of the target, a black pedestal fan. The game said FOUND IT . The log showed why. The model described my laptop photo as "Black pedestal fan with metal grille", the target's own words. My prompt told it what the target was before it looked at my photo, and a small model, told what to expect, sees it. The check now happens in three steps: On my 24 test pairs it scores the same as before 10/10 real finds, 6/7 look-alikes, 7/7 unrelated , and it rejects the laptop. A wrong photo now gets its NOT QUITE in about 3 seconds. The first version only worked one way: stand somewhere, take the photo, wait, hunt, all with the phone out and a laptop on the same network. I couldn't test that in the middle of the street, and not everyone can, or should, walk around staring at a phone. So now the photo can come from your library too, hunts stay open for days, and the close-up can be a five-second photo on your way past, checked whenever you like. You can still play it all on the spot. That turned the slow model and the laptop at home into non-issues: you set up a hunt at home, look on your normal route, and check at home again. The screen part happens indoors. The street part doesn't need a screen. 276 tests, including the engine replayed against recorded real Gemma replies , so the pipeline is tested on what the model actually says, deterministically, without a GPU. Accessibility checks with axe on every screen. Best Use of Gemma. Gemma 4 E4B does all the seeing: it proposes targets with native box 2d boxes, verifies each one on its own crop, describes the player's close-up without knowing the answer, and compares the two images. All locally, on a laptop.