FloraTrail AI A developer built FloraTrail AI, a fully offline, screenless voice assistant for hikers that runs on a Raspberry Pi 5 or pocketable Linux/Android device. The system combines whisper.cpp for on-device speech transcription, a V4L2 camera capture script, and a local vision-language model such as LLaVA-13B or Moondream2 served through Ollama, with Piper TTS reading naturalist answers into an earpiece. The demo shows the stack identifying a mushroom and answering foraging-safety questions in a zero-connectivity area of the Pacific Northwest. This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass https://dev.to/challenges/hacktoberfest-week1-2026-10-05 Meet FloraTrail AI, a fully offline, voice-assisted wilderness companion designed specifically to help you disconnect from screens and immerse yourself in nature. While most AI tools keep us tethered to our desks or glued to our smartphones, FloraTrail flips the paradigm. It is an entirely screenless, audio-first local agent that runs on your backpack's Raspberry Pi 5 or any pocketable Linux/Android device . You snap a photo using a connected camera, tap a hardware button on your jacket lapel to ask a question e.g., "What kind of bird is making that call, and what does it eat?" , and FloraTrail whispers the answer directly into your earpiece. It actively encourages users to "touch grass" by removing the UI layer entirely. You stay immersed in the sights, sounds, and smells of the forest, receiving expert naturalist insights without ever breaking eye contact with the trail. It is custom-built for hikers, birdwatchers, foragers, and off-grid adventurers who venture far beyond the reach of cell towers. In the demo, I am hiking through the deep woods of the Pacific Northwest—a verified zero-connectivity dead zone. I spot a strange, brightly colored mushroom. I simply point my lapel camera at it, tap my Bluetooth mic, and ask, "Is this safe to forage?" Because the entire stack is local, the inference happens on my device in seconds. The local vision model analyzes the frame, the text agent formats the safety warning, and the local TTS engine reads it back to me. No loading spinners, no "No Internet Connection" errors—just instant, offline augmented reality for nature. import base64 import ollama import subprocess def analyze trail sighting image path: str, transcribed audio: str - str: """ Sends the captured image and transcribed voice prompt to the local Vision-Language Model. """ 1. Encode the locally captured image with open image path, "rb" as image file: encoded image = base64.b64encode image file.read .decode 'utf-8' system prompt = "You are an expert botanist, mycologist, and ornithologist. " "Keep your response concise, conversational, and under 3 sentences, " "as it will be read via Text-to-Speech to a hiker. Emphasize safety." print " Running local inference on LLaVA/Gemma..." 2. Invoke local vision model via Ollama response = ollama.chat model='llava:13b', Swappable with other open weights like Moondream2 messages= { 'role': 'user', 'content': f"{system prompt}\n\nHiker's question: '{transcribed audio}'", 'images': encoded image } insight = response 'message' 'content' 3. Pass the response to local TTS Piper print f" AI Naturalist says: {insight}" subprocess.run "piper", "--model", "en US-lessac-medium", "--output file", "/tmp/output.wav" , input=insight.encode 'utf-8' subprocess.run "aplay", "/tmp/output.wav" return insight To make FloraTrail a reality without relying on cloud APIs, I had to architect a robust, lightweight local agent harness. Here’s the step-by-step breakdown of the open-source stack: Audio Ingestion Whisper.cpp : When the user presses the push-to-talk button, audio is recorded and instantly transcribed using whisper.cpp running a quantized tiny.en model. It takes milliseconds and uses less than 100MB of RAM. Image Capture: A bash script triggers the camera module via V4L2 Video4Linux to capture a high-res frame of whatever the user is looking at. Local Agent Harness Python/LangChain : A custom Python loop acts as the orchestrator. It bundles the transcribed voice prompt and the base64-encoded image. Local AI Inference Engine Ollama / llama.cpp : The bundle is sent to an open-weight vision-language model LLaVA-13B or Moondream2 for ultra-low resource devices running via Ollama. The model interprets the visual data in the context of the user's audio question. Text Formatting Gemma-2 : Because VLM outputs can sometimes be overly technical or contain markdown which ruins text-to-speech , the raw output is passed through a lightweight gemma-2-2b instruction model to re-write the text into a conversational, TTS-friendly script. Audio Output Piper TTS : Finally, Piper a fast, local neural text-to-speech engine synthesizes the text into a human-sounding voice and plays it back through the user's headphones. Building this project underscored exactly why open-source AI is the only viable future for ubiquitous computing: Off-Grid Reliability: The best parts of nature do not have 5G. Cloud APIs fail completely on remote hiking trails. By running quantized open-weight models on edge devices, AI becomes a reliable utility everywhere on Earth, not just in urban centers. Privacy & Zero-Data Leakage: When you are hiking, you are in your sanctuary. Centralized cloud services harvest location metadata, voice prints, and personal photo streams. Local open-source AI ensures absolute data sovereignty—what happens on the trail stays on the trail. Zero Operating Cost: Cloud AI models charge per token. When exploring nature, you might want to query hundreds of plants, bugs, and tracks in a single weekend. Local inference provides infinite scaling at zero marginal cost. Customizability: Open weights allow for incredible hyper-localization. For a future update, I plan to use LoRA to fine-tune the vision model specifically on the indigenous flora of my local state parks, achieving higher accuracy for my specific region than a generalized commercial model ever could.