This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
Most apps are designed to keep you in them. Step Outside is designed to get rid of you.
It's a small web app that waits while you work, then taps you on the shoulder: "You've been looking down for a while." It gives you one challenge, like find something alive growing somewhere unusual. You step away from the screen, take one photo, and an AI replies with a single calm sentence about what you found. Then it says the thing most apps never say:
"Enough. Put your phone away."
It doesn't just show the line on screen. It says it out loud too, so the last thing you hear is the app telling you to stop, not another notification asking you to stay.
The session closes, and it stays closed. The server refuses any further requests for it, and the only way back in is a deliberate "Start again" tap. There are no streaks, no feed and no "one more round".
You choose how long it waits before nudging you, from 15 seconds for demos to four hours. It's for anyone who has noticed that their eyes have been 40 centimeters from a screen since breakfast, and who would like a nudge without a notification hole to fall into.
An app that gets you OFF your screen: nudge β challenge β one photo β "Enough. Put your phone away."
Create a .env file in the project root, next to main.py. Add the keys for the integrations you want to use:
GEMINI_API_KEY=your_gemini_api_key
SERPAPI_KEY=your_serpapi_key
MONGODB_URI=your_mongodb_connection_string
SENTRY_DSN=your_sentry_dsn
pip install -r requirements.txt
uvicorn main:app --reload
Reminder time: pick it on the start screen (in minutes). For a quick demo, enter 0.25 (= 15 seconds).
Session flow: idle β challenged β responded β ended. After "Enough.", a Start again button returns to the start screen.
| File | What it does |
|---|---|
main.py |
Web server + API routes |
ai.py |
Gemma: makes challenges, reacts to photos |
db.py |
MongoDB Atlas storage |
extras.py |
SerpApi info about what you found |
static/ |
The web page (HTML, CSS, JS) |
Push to GitHub β Renderβ¦
Gemma is the heart of the app. It does two jobs: it writes the challenge, and it looks at the photo and describes what it sees. The prompt asks for two lines, a SUBJECT and a one-sentence REPLY, which is easy to parse and easy to keep short. I use Gemma 4 through the Gemini API, with gemma-4-26b-a4b-it as the default and gemma-4-31b-it as a backup. Photos are shrunk to 768px first, which keeps uploads fast and is plenty to recognise grass or a cloud.
MongoDB Atlas stores each session: its state, the challenge, the reply, the subject and how long the person was on the page. The photo itself is never saved. Challenges are saved too, so Gemma is told what it has already asked and avoids repeating itself. This is a tiny, simple memory.
SerpApi takes the SUBJECT Gemma identified and fetches a short fact about it, so a photo of a weed can become "here's what it's called".
Sentry traces the AI steps (make_challenge, react_to_photo), so I could see where time was going and when something failed.
Render hosts it.
Browser speech reads the ending out loud. It uses the browser's built-in voice, so it needs no extra service or API key. Right after Gemma's reply, the page says "Nice. That's enough. Put your phone away.", then the session closes.
The session is a small state machine: idle β challenged β responded β ended. Each endpoint accepts only the correct previous state, and ended returns an HTTP 410 with "Enough. Put your phone away." That one rule is what makes it an anti-engagement app, and it takes about five lines.
A few things went differently than planned:
new Notification(). try/catch and showing an error message fixed it. Phones also timers, so the app now checks the real clock.
1. I could swap the model by changing one name.
Partway through the build, the Gemma model I started with stopped working because Gemma 4 had replaced it. Fixing it meant changing a model name in an environment variable. Gemma comes in more than one size, so the app tries the lighter Gemma 4 model first and falls back to the larger one. This paid off when Sentry caught a "Deadline expired" timeout from a slow call: the retry and fallback absorbed it, and the person using the app never noticed. I could trade speed against quality without rewriting anything around the model.
2. One open model handles both jobs, and the app only needs a sentence from it.
Gemma 4 reads text and images, so the same model writes the challenge and then looks at the photo. Step Outside is built around saying less, so it doesn't need a giant model. It needs a small one that can be told exactly how to behave: one calm sentence, no judging, no questions, nothing that invites another round. The "Enough. Put your phone away." ending lives in my own code, not in the model, so the model supplies the observation and my app decides when the conversation is over.