PIP: a desktop penguin that loves you but hates it when you stay on the desktop for too long A developer built Pip, a Windows desktop penguin that tracks screen time and locks the user's screen after 65 minutes until they complete an offline quest verified by photo. The system runs entirely locally in a single Python file using Gemma 3 4B via Ollama for personality, answer judging and vision labeling, plus Vosk for speech-to-text, OpenCV for face and posture checks, and pyttsx3 for voice. The developer reports the small 4B vision model proved harder to fool than expected, correctly distinguishing leaves from bark and rejecting reused photos, and concludes that code should decide while the model talks. This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass https://dev.to/challenges/hacktoberfest-week1-2026-10-05 Pip is a small, chronically online penguin who lives on your Windows desktop. His life mission is to get you off the screen. He tracks how long you've really been at the computer: Quests Proof, so you can't just click "done" The rest of Pip I took Pip outside and tried to fool him. It's harder than it sounds. The leaf quest Pip told me to go for a walk and bring him a leaf. The old-photo test I also tried showing him an old picture instead of going out. He rejected it, because he remembers the pictures he's already accepted and won't take the same one twice. Trying to trick the vision model What surprised me: I thought a 4B model running locally would be easy to fool, especially with look-alikes. It wasn't. It told the difference between a leaf and bark, a pebble and a shell, and grass and a pile of leaves. The result: the only way to get my screen back was to get up, go outside and bring back the right thing. Pip is a small, chronically online penguin who lives on your Windows desktop. His one mission is to get you off the screen. He tracks how long you've really been at your computer. At 30 minutes he raises a hand and nudges you. At 60 minutes he turns red and starts steaming. Five minutes later he locks your screen until you pick a quest and prove you did it. Everything runs locally: the language model, the vision model, speech recognition and the camera checks. Setup instructions are in the README. You need Windows, Python, and Ollama https://ollama.com with gemma3:4b . Everything runs on my own Windows PC in a single Python file, with no paid API. | Piece | What it does | |---|---| | Gemma 3 4B via Ollama , open-weight | Pip's personality and replies, the answer judge, and his diary entries | | Gemma 3 vision same model | Looks at the photo you hold up and labels it leaf, flower, pebble, grass, cup... | | Vosk | Offline speech-to-text, so I can just talk to Pip | | OpenCV | Checks whether a face is in view the "are you really away" test , blink rate and posture, all in memory | | pyttsx3 | Offline voice with an optional neural voice if edge-tts is installed | | Open-Meteo | Free weather, no API key | | Tkinter | Draws the see-through penguin and the whole UI | The main lesson from using a small local model: code decides, the model talks. A 4B model is a great character but a bad referee. So the timers, the screen lock, sound effects and the weather sanity check are plain code. Gemma only writes Pip's lines and judges answers inside strict JSON prompts. Verification is layered too: the photo label, then a duplicate-photo check, then screen and printed-picture rejection. If the judge ever crashes, it doesn't punish the user. Pip watches me through a webcam, listens through a mic, knows my to-do list and writes a diary about my day. I would never want to stream that to a server I don't control. Best Use of Gemma: Pip's brain runs on Gemma 3 4B locally through Ollama. Gemma writes Pip's personality and replies, judges whether your answer to his question really holds up, writes his diary entries, and using its vision ability checks the photo you hold up to the camera. If the photo check is sloppy, I can swap in a bigger Gemma by changing one line.