{"slug": "trail-tutor", "title": "TRAIL TUTOR", "summary": "A developer built TrailTutor AI, a FastAPI-based outdoor learning companion that generates short, structured outdoor missions for learners using Google's DiffusionGemma-26B model via NVIDIA's hosted API. The developer also LoRA fine-tuned Qwen3.5-4B with Tinker on a small reproducible dataset, improving held-out structured-output scores from 95% to 100% and cutting average latency by about 13.4% (3.355s to 2.907s). The project documents a corrected v1.1 evaluation after an initial parser bug inflated the reported improvement.", "body_md": "*This is a submission for the [Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass](https://dev.to/challenges/hacktoberfest-week1-2026-10-05)*\n\nTrailTutor AI is an outdoor learning companion designed around a simple idea:\n\n**AI should sometimes help us leave the screen, not stay on it.**\n\nA learner chooses:\n\nTrailTutor then generates one short outdoor mission with:\n\nThe learner reads the mission, puts the device away, completes the activity outdoors, then comes back only to reflect.\n\nThat is why I designed TrailTutor around the principle:\n\nThe screen should be the shortest part of the experience.\n\nRENDER APP\n\n[https://trailtutor-ai.onrender.com/](https://trailtutor-ai.onrender.com/)\n\nLocal host:[http://127.0.0.1:8000/](http://127.0.0.1:8000/)\n\n[https://github.com/rajab-rajab/TrailTutor-AI](https://github.com/rajab-rajab/TrailTutor-AI)\n\nThe public application is a lightweight FastAPI service deployed on Render.\n\nFor the live Gemma path, TrailTutor uses:\n\n`google/diffusiongemma-26b-a4b-it`\n\nthrough NVIDIA's hosted API.\n\nA typical request contains:\n\n```\n{\n  \"age\": 13,\n  \"environment\": \"school playground\",\n  \"topic\": \"plants\",\n  \"duration_minutes\": 10\n}\n```\n\nThe response follows a small structured format:\n\n```\n{\n  \"mission\": \"...\",\n  \"observe\": \"...\",\n  \"questions\": [\"...\", \"...\"],\n  \"safety\": \"...\",\n  \"reflection\": \"...\"\n}\n```\n\nThis structure is intentionally restrictive. TrailTutor is not supposed to become another long conversation. Its job is to produce a useful mission quickly and then get out of the learner's way.\n\nI wanted the core AI component to use an open-weight model rather than treat the model as a completely closed black box.\n\nGemma gave me a strong foundation for generating short, structured educational activities while keeping the architecture flexible enough to swap or self-host models later.\n\nThe live TrailTutor application uses DiffusionGemma for mission generation.\n\nI also wanted to test whether a general open-weight model could be specialized specifically for TrailTutor's mission format.\n\nI used Tinker to LoRA fine-tune:\n\n`Qwen/Qwen3.5-4B`\n\nThe training setup was deliberately small and reproducible:\n\nThe dataset covers:\n\nThe held-out prompts have no exact input overlap with the training examples.\n\nI evaluated the untuned and tuned versions of the same Qwen3.5-4B model on the 20 held-out prompts.\n\nThe final corrected evaluation produced:\n\n| Metric | Untuned | Tinker-tuned | \n|---|---|---|\n| Overall structured-output score | 95% | 100% | \n| Valid JSON | 19/20 | 20/20 | \n| All required fields | 19/20 | 20/20 | \n| Exactly two questions | 19/20 | 20/20 | \n| Outdoor action present | 19/20 | 20/20 | \n| Safety present | 19/20 | 20/20 | \n| Reflection present | 19/20 | 20/20 | \n| Mission under 70 words | 19/20 | 20/20 | \n| Average latency | 3.355 s | 2.907 s | \n\nThat is a **5 percentage-point improvement** in the held-out structural score.\n\nAverage latency also decreased by about **13.4%**.\n\nMy first evaluation reported a larger improvement: 60% to 80%.\n\nThat result turned out to be misleading.\n\nSome model outputs contained two consecutive valid JSON objects. My original parser took everything from the first opening brace to the last closing brace, then tried to decode the whole string as one JSON object.\n\nThat caused valid generations to be scored as failures with errors such as:\n\n`Extra data`\n\nInstead of hiding that mistake, I kept the original result in the repository and added a corrected v1.1 evaluation that parses the first complete JSON object while preserving the raw output.\n\nThe corrected result is the one I report here:\n\n**95% → 100%**\n\nI think keeping both versions is important because reproducible AI evaluation also means documenting when the evaluator itself was wrong.\n\nTrailTutor is deployed publicly on Render using the project's Dockerfile.\n\nRender hosts:\n\nThe computationally heavy Gemma inference happens through the model provider, so the web service itself stays lightweight.\n\nThis made it straightforward to turn the local prototype into a publicly accessible project that judges and users can try immediately.\n\n```\nUser\n |\n v\nTrailTutor web interface\n |\n v\nFastAPI application on Render\n |\n +----------------------------+\n |                            |\n v                            v\nDiffusionGemma             Tinker experiment\nvia NVIDIA API             Qwen3.5-4B\n                              |\n                              +--> untuned baseline\n                              |\n                              +--> LoRA-tuned model\n```\n\nTrailTutor benefits from open models because I can do more than simply send text to an opaque endpoint.\n\nI can:\n\nThe Tinker experiment is a good example.\n\nInstead of claiming that fine-tuning helped, I could actually train an open-weight model, evaluate it against its own untuned baseline, inspect failures, discover a bug in my evaluator, and publish the corrected evidence.\n\nThat kind of experimentation is much easier when the model ecosystem is open enough to adapt.\n\nThe biggest surprise was not the training result.\n\nIt was the evaluation bug.\n\nSeveral outputs I initially counted as complete failures were actually good TrailTutor missions repeated twice.\n\nThat reminded me that model evaluation is not only about testing the model. The evaluator also needs to be tested.\n\nIt changed the final result from an apparent 20-point improvement to a more defensible 5-point improvement.\n\nI prefer the smaller number because I can explain exactly where it came from.\n\nI would like to continue in three directions:\n\nI am entering TrailTutor AI for:\n\nIt is also automatically eligible for the overall **Touch Grass** challenge.\n\nMany AI applications are designed to increase engagement time.\n\nTrailTutor deliberately tries to do the opposite.\n\nIt uses AI to give a learner one useful reason to close the screen, step outside, pay attention to the physical world, and come back with something they noticed for themselves.", "url": "https://wpnews.pro/news/trail-tutor", "canonical_source": "https://dev.to/rajab_baig_a3929cefc3758b/trail-tutor-el7", "published_at": "2026-10-09 23:33:35+00:00", "updated_at": "2026-10-09 23:58:13.385858+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-tools", "ai-products"], "entities": ["TrailTutor AI", "Google", "DiffusionGemma-26B", "NVIDIA", "Tinker", "Qwen3.5-4B", "FastAPI", "Render"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/trail-tutor", "markdown": "https://wpnews.pro/news/trail-tutor.md", "text": "https://wpnews.pro/news/trail-tutor.txt", "jsonld": "https://wpnews.pro/news/trail-tutor.jsonld"}}