#
What I Built
I built Trail Council, a panel of four small AI agents that runs on my laptop with no internet and decides whether I should go on a hike today, and which route to take.
Four agents each look at the plan from one angle: Weather, Route, Gear and Safety. They debate, and a final "lead" agent gives one verdict: Go, Go with a shorter route, Gather more info, or Stay home. The output is a one-page printable plan with the route, turnaround time, gear checklist and the reasoning. I print it, close the laptop, and walk out the door.
The screen is the shortest part of the trip. I use it for about two minutes before I leave, and I can regenerate the plan at the trailhead with zero signal.
It's for anyone who plans outings in places with bad connectivity: hikers, trail runners, and weekend trekkers.
#
Code
I reused the architecture of an earlier project of mine (AI Boardroom, a multi-agent business decision simulator). Trail Council is a new repository with new agents, new data, new prompts and a new output format. Only the staged pipeline idea carries over.
#
How I Built It
Stack
Model: Qwen2.5-3B-Instruct, an open-weight model, loaded in 4-bit (NF4) with PyTorch andtransformers #
Hardware: a laptop with an RTX 3050 (4GB VRAM). No cloud GPU. #
Backend: FastAPI + Uvicorn #
Retrieval:all-MiniLM-L6-v2 embeddings + ChromaDB over a small local knowledge base ([YOUR DATA: e.g. my own route notes, sunrise/sunset table, gear checklists, local trail descriptions]) #
Frontend: [HTML/JS page or Streamlit] The four-stage pipeline
Agent analysis: Weather, Route, Gear and Safety each read the same trip request plus their own retrieved notes and return structured JSON. 2. Debate: Agents flag disagreements, for example Safety objecting that the long loop finishes after sunset. 3. Lead decision: One agent weighs the analyses and picks a verdict. 4. Validation: A check confirms the verdict is one of the four allowed categories and that the next step actually follows from it. Otherwise it retries.
What was hard
A 3B model on 4GB of VRAM drops instructions, cuts off mid-answer and drifts from the format. Three things fixed most of it:
- JSON-schema-constrained outputs
- Short, role-specific prompts with grounding hints
- Compact "cards" passed between stages instead of full text, plus a retry when output is truncated
[ADD 1-2 REAL NUMBERS FROM YOUR MACHINE: time to generate a full plan, VRAM used, how often validation had to retry.]
#
Field Test
I took it to [PLACE] on [DATE]. #
What worked: [e.g. it flagged that my planned loop would end after sunset and suggested a shorter one] #
What didn't: [e.g. the Weather agent was confidently wrong about X, because my data was stale] #
What I changed on the trail: [e.g. regenerated the plan at the trailhead with no signal, and it worked] #
Did I follow its advice? [yes/no, and what happened]
#
Why Does Open Innovation Matter?
It worked with no signal. The model, the embeddings and the vector database all live on my laptop. A hosted-API version of this would be useless at exactly the place it's needed, a trailhead with no coverage. #
My plans stay mine. Where I go, when I'll be alone, and when I'll be back are sensitive details. Nothing is sent to a server I don't control. #
It costs nothing to run. No API key and no per-request bill. I regenerated plans [N] times while testing. #
I could tune the behavior. I rewrote agent roles, swapped the model between sizes, and changed the validation rules. I couldn't have done that with a closed black box. #
The honest tradeoff: A 3B model is less capable than a frontier model. I had to design around its limits with schemas, short prompts and validation. For this task I think that's a fair trade, but it is a trade.
#
Prize Categories
Best Use of Render: [ONLY IF you really host the demo frontend on Render; say what runs there and what runs locally]