Trail Council: Four Tiny AI Agents on My Laptop That Decide If I Should Go Hiking A developer built Trail Council, a fully offline multi-agent system that runs on a laptop with an RTX 3050 (4GB VRAM) and decides whether to go hiking and which route to take. Four agents — Weather, Route, Gear and Safety — analyze a trip request using Qwen2.5-3B-Instruct in 4-bit quantization plus local MiniLM embeddings in ChromaDB, debate disagreements, and pass compact cards to a lead agent that issues one of four verdicts, validated against an allowed set with retries. The project reuses the staged pipeline architecture from the developer's earlier AI Boardroom simulator and outputs a one-page printable plan that can be regenerated at a trailhead with no signal. What I Built I built Trail Council , a panel of four small AI agents that runs on my laptop with no internet and decides whether I should go on a hike today, and which route to take. Four agents each look at the plan from one angle: Weather , Route , Gear and Safety . They debate, and a final "lead" agent gives one verdict: Go , Go with a shorter route , Gather more info , or Stay home . The output is a one-page printable plan with the route, turnaround time, gear checklist and the reasoning. I print it, close the laptop, and walk out the door. The screen is the shortest part of the trip. I use it for about two minutes before I leave, and I can regenerate the plan at the trailhead with zero signal. It's for anyone who plans outings in places with bad connectivity: hikers, trail runners, and weekend trekkers. Code I reused the architecture of an earlier project of mine AI Boardroom https://github.com/krushnakodgirwar/AI-Boardroom , a multi-agent business decision simulator . Trail Council is a new repository with new agents, new data, new prompts and a new output format. Only the staged pipeline idea carries over. How I Built It Stack - Model: Qwen2.5-3B-Instruct, an open-weight model, loaded in 4-bit NF4 with PyTorch and transformers - Hardware: a laptop with an RTX 3050 4GB VRAM . No cloud GPU. - Backend: FastAPI + Uvicorn - Retrieval: all-MiniLM-L6-v2 embeddings + ChromaDB over a small local knowledge base YOUR DATA: e.g. my own route notes, sunrise/sunset table, gear checklists, local trail descriptions - Frontend: HTML/JS page or Streamlit The four-stage pipeline 1. Agent analysis: Weather, Route, Gear and Safety each read the same trip request plus their own retrieved notes and return structured JSON. 2. Debate: Agents flag disagreements, for example Safety objecting that the long loop finishes after sunset. 3. Lead decision: One agent weighs the analyses and picks a verdict. 4. Validation: A check confirms the verdict is one of the four allowed categories and that the next step actually follows from it. Otherwise it retries. What was hard A 3B model on 4GB of VRAM drops instructions, cuts off mid-answer and drifts from the format. Three things fixed most of it: - JSON-schema-constrained outputs - Short, role-specific prompts with grounding hints - Compact "cards" passed between stages instead of full text, plus a retry when output is truncated ADD 1-2 REAL NUMBERS FROM YOUR MACHINE: time to generate a full plan, VRAM used, how often validation had to retry. Field Test I took it to PLACE on DATE . - What worked: e.g. it flagged that my planned loop would end after sunset and suggested a shorter one - What didn't: e.g. the Weather agent was confidently wrong about X, because my data was stale - What I changed on the trail: e.g. regenerated the plan at the trailhead with no signal, and it worked - Did I follow its advice? yes/no, and what happened Why Does Open Innovation Matter? - It worked with no signal. The model, the embeddings and the vector database all live on my laptop. A hosted-API version of this would be useless at exactly the place it's needed, a trailhead with no coverage. - My plans stay mine. Where I go, when I'll be alone, and when I'll be back are sensitive details. Nothing is sent to a server I don't control. - It costs nothing to run. No API key and no per-request bill. I regenerated plans N times while testing. - I could tune the behavior. I rewrote agent roles, swapped the model between sizes, and changed the validation rules. I couldn't have done that with a closed black box. - The honest tradeoff: A 3B model is less capable than a frontier model. I had to design around its limits with schemas, short prompts and validation. For this task I think that's a fair trade, but it is a trade. Prize Categories - Best Use of Render: ONLY IF you really host the demo frontend on Render; say what runs there and what runs locally