This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
I built Hue Hunt, a daily colour hunt for a group of friends.
Every morning the group gets a colour and a bonus twist, for example "Yellow. Bonus: something that moves." During the day everyone goes out and photographs what they find. A yellow car, a bin lid, a sunflower. Nobody sees anyone else's photos until the evening reveal, when all of them appear side by side with points, a winner and a short, playful recap. A season runs for a week, with a leaderboard at the end.
The judge is an AI model. It looks at each photo and answers a few questions: is the thing in the photo really yellow, does it move, what is it, and is it a photo of a screen or a poster rather than a real thing. It also writes a one-line comment on every photo, which turned out to be the part people react to most. "That tiny yellow logo is a bit lost in all that gorgeous red!" is a fair verdict on a red Ferrari submitted on a yellow day.
The screen is the shortest part of the game. You read the colour in the morning, which takes five seconds. Taking a photo takes another five. Then you look at the reveal in the evening. Everything in between happens outside, and the game is a reason to look at things you walk past every day.
It's for a group of friends, the kind who already have a chat group. There are no accounts. You share a link, people type a name, and their phone remembers them. The person who creates the group is the host and can set the reveal time, press "reveal now", or start a new season.
Hue Hunt runs on a small private server for my friends and me, so there's no public link yet. Running it yourself takes five minutes, and the steps are in the next section.
Here is a demo season, screenshotted from the real app. I seeded it with a folder of photos and let the real judge score them, so every point and comment below is Gemma's.
Day 3: orange, "something with a face". The cat has been judged: 15 points, bonus earned, and a comment.
The reveal for the yellow day. Maja wins with a taxi, a raincoat and a sunflower. Every photo has its points and the judge's comment. Tom's sunlit grass got 0: "That golden sunlight is beautiful, but we're hunting for actual yellow objects!"
The season leaderboard after two days.
The invite page. No account, no password. The small print says photos go to a third-party API and that location data is stripped.
The share links carry preview cards, so a reveal unfurls in WhatsApp or iMessage as a collage of the day's best photos.
A daily colour hunt for a group of friends, judged by an open-weight vision model.
Every morning the group gets a colour and a bonus twist, for example "Yellow. Bonus: something alive." During the day everyone photographs what they find outside, straight from the phone's camera in the browser. Nobody sees anyone else's photos until the evening reveal, when all the photos appear side by side with scores, a winner and a short, playful recap. Seasons run a week with a leaderboard at the end.
Built for the DEV Hacktoberfest 2026 "Touch Grass" challenge. The screen is the shortest part of the game: a few seconds to take a photo, one look at the reveal in the evening.
The README has the full setup. In short, you need Node.js 24 and pnpm, and then:
git clone https://github.com/vytasgavelis/hue-hunt.git
cd hue-hunt
cp .env.example .env # add a free Gemini API key from Google AI Studio
pnpm i
pnpm db:migrate
pnpm dev
Then open http://localhost:5201. If you want to look at a few files, these are the interesting ones:
apps/server/src/ai/tasks.ts has the four prompts: judge a photo, invent the day's challenge, write the evening recap, write the season recap.apps/server/src/ai/client.ts is the model client. It asks for JSON, validates it, retries once, and logs every call.packages/shared/src/scoring.ts turns the judge's answers into points. No AI in this file.apps/server/src/scripts/eval-judge.ts is the judging eval, which is how I chose the model and tuned the prompt.apps/server/src/lib/images.ts strips the metadata from every photo before it is stored.
Hue Hunt is a TypeScript app: a React web app, a Hono API, a SQLite database and a small job queue, all in one Node process on a $6 server. The AI side is one open-weight model, Gemma 4 26B-A4B from Google, called through the Vercel AI SDK. I call it on the Gemini API's free tier, so running Hue Hunt needs no GPU and no paid account, but the model is a published set of weights and nothing in the code depends on where it runs.
Gemma does four things, and they're all single prompt-and-reply calls. No agents, no tools, no streaming.
Judging a photo is the only task that sees images. The prompt gives the colour and the twist, and asks for exactly this:
{
"colour_match": 0.0, // number 0 to 1 on the scale above: how clearly the thing photographed is "yellow"
"bonus_match": false, // true only if the yellow thing in the photo itself satisfies the bonus twist "something that moves"
"subject": "...", // short noun phrase naming the main subject, e.g. "yellow taxi"
"category": "...", // one of: vehicle, plant, animal, building, sign, clothing, food, object, other
"is_screen_or_print": false, // true if this is a photo of a screen, a poster, a printed image or a drawing rather than a real thing
"comment": "..." // one playful line about the photo, at most 15 words, friendly never mean
}
This is the part a pixel histogram can't do. "Yellow and alive" is a question about what the thing is, not what colour the pixels are. So is "it's just a tiny sign in the corner of a grey street", and so is "that's a photo of a laptop screen".
Inventing the challenge runs at 7 am in the group's timezone. The model gets the colours and twists already used this season and returns a new colour, a twist and a tagline. It likes "terracotta" and "cobalt blue" more than I expected.
The recaps never see the photos. The evening recap gets the judge's notes for each player, the subject, points and comment of every photo, and writes three to five sentences naming the winner and the funniest find. The season recap does the same across the week.
The model answers questions about a photo. Points are worked out by plain code that I can test:
export function photoPoints(j: Judged): number {
if (j.is_screen_or_print) return SCORING.screenOrPrintPoints;
const bonus = j.bonus_match && j.colour_match >= SCORING.bonusMinColourMatch;
return Math.round(j.colour_match * SCORING.colourPoints) + (bonus ? SCORING.bonusPoints : 0);
}
Colour match times 10, plus 5 for the twist, 0 for screens. Your best three photos count for the day, and a second photo in the same category counts half, so three photos of three yellow cars don't beat one car, one flower and one bin. The season adds 3 points for each daily win.
The bonus only counts when the colour match is at least 0.5. That rule came from the eval. On a yellow day with "something alive", the model correctly called a green forest "alive", and without the rule it would have scored 5 points for a photo with no yellow in it.
Gemma's native JSON schema mode let stray tokens into values in my last project, so this time I never turned it on. The prompt shows a filled-in template of the reply, with a comment on each field, and the server validates the reply with a Zod schema that the web app shares. The schema is forgiving where it can be: a colour match of "0.8" as a string is coerced to a number and clamped to 0 to 1, a category the model made up becomes "other". If the reply still doesn't validate, the call is retried once with a reminder. Of the 57 model calls logged while I built and tested it, none needed the retry.
Setting Gemma's thinking to "minimal" matters a lot here. Judging a photo takes 3 to 5 seconds. The recaps are slower, around 30 seconds, because they are longer and I let the model be more creative.
I didn't want to guess, so I wrote pnpm eval:judge. It runs the judge over a folder of photos for a given colour and twist and prints a table. The photos are named for what I expect, so the table reads like a test report.
I ran it on nine photos for "yellow, something alive": a taxi, a raincoat, a sunflower, bananas, a duckling, a green forest, a red Ferrari, a photo of a smartphone screen, and a yellow mosaic on a wall.
| Gemma 4 26B-A4B | Gemma 4 31B | |
|---|---|---|
| Photos judged | 9 / 9 | 5 / 9 (4 timeouts at 60 s) |
| Judgements I agreed with | 9 | 5 |
| Median time per photo | 3.8 s | 40.8 s |
Both models judged the photos the way I would have. The 31B is dense, the 26B is a mixture-of-experts model with about 4 billion active parameters, and it's ten times faster. That's the difference between a comment appearing while you're still looking at your photo and a comment appearing after you've put the phone away, so the 26B is the default.
The first prompt was two sentences. Then I ran it on a second set of photos, the kind of thing you'd actually find on a walk, and the results were uneven:
| Photo | First prompt | Final prompt |
|---|---|---|
| Orange cat, "something alive" | 0.70 + bonus = 12 pts | 0.30 = 3 pts |
| Orange traffic cone | 0.10 | 0.30 |
| Yellow lid on a green bin | 0.40 | 0.80 |
| Sunlight on grass | 0.40 + bonus | 0.00 |
| Field of sunflowers | 1.00 + bonus | 1.00 + bonus |
An orange cat scoring 12 points on a yellow day was the problem. The model knew it was orange, its comment said so, but it gave 0.7 anyway. The cone, which is just as orange, got 0.1. And the bin lid, which is unmistakably yellow and the obvious point of the photo, got 0.4 because the bin around it is green.
The fix was to tell the judge what I actually meant. The final prompt judges the hue strictly and the size leniently: any shade of the colour counts fully, a neighbouring hue caps at 0.3, a coloured part of a bigger object counts if it's clearly the point of the photo, and light cast by the sun doesn't count. It also gives an anchored scale, 1.0, 0.8, 0.5, 0.3, 0, with a sentence for each.
One lesson from the iteration. My first version of that rule said "orange, gold or ginger is not yellow". Gemma took "gold" literally and dropped a golden-orange sunflower from 1.0 to 0.3. The rule became "lemon, mustard and golden yellow are all yellow" instead, and the sunflower came back. With a model this fast, the loop of change the prompt, run the eval, read the table takes a minute.
Photos are sent to a third-party API for judging. The join page says so. Before a photo is stored, the server decodes it, rotates it the right way up, resizes it to 1280 px and writes a fresh JPEG. All metadata is gone, including the GPS location phones embed in every shot. There's a test that uploads a photo with GPS coordinates and asserts they are not in the file on disk. Photos are only visible to the group, and other people's photos are hidden until the reveal.
I wrote a spec first, then built it with Claude Code as my coding assistant over two days. The spec was the useful part. It settled the scoring rules, the data model and the job queue before any code existed, and it meant the first version came out working end to end, with 40 tests. I then spent my time on the eval and the prompt, which is where the quality of the game actually lives.
The judge can be argued with. A game with an AI referee only works if people trust the referee. Everything the judge is told is in one file, the prompt is plain text, and anyone can read it, run the eval on their own photos, and see the table. When my friends disagree with a score, I can show them the rule, and if they are right I change it. A closed, hosted judging feature would be a black box with a score coming out of it.
The model I tested is the model I ship. Gemma 4 26B-A4B is a set of published weights. The eval results above stay true, because the model can't be silently updated or retired under me. With a closed API, the provider can swap the model next month and the orange cat might be worth 12 points again.
It can run anywhere, including at home. Today the calls go to Google's free tier because it was the easiest. The provider is one environment variable, and the code has a second provider for any OpenAI-compatible server. A group that doesn't want its photos leaving the house can run the same model on a machine at home with Ollama or vLLM and get the same judge. That's a real option for a game where people photograph their street, their garden and their kids' bikes.
Cost makes the game possible at all. A group of five taking five photos a day is 25 model calls a day, plus a challenge and a recap. A small open model on a free tier makes that free. A frontier vision model at API prices would make a daily game for friends an expense nobody would pay.
The frameworks are open too. The Vercel AI SDK is a library inside my app. When I wanted the image passed one way rather than another, or the thinking level set, I read the source instead of waiting for a feature.
Gemma. Hue Hunt runs on Gemma 4 26B-A4B for every model call: judging photos, inventing the daily challenge and writing the recaps. It was chosen over Gemma 4 31B by the eval above.
Hue Hunt is a first version, and it has gaps I know about:
What I want to build next: