This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
TrailTutor AI is an outdoor learning companion designed around a simple idea:
AI should sometimes help us leave the screen, not stay on it.
A learner chooses:
TrailTutor then generates one short outdoor mission with:
The learner reads the mission, puts the device away, completes the activity outdoors, then comes back only to reflect.
That is why I designed TrailTutor around the principle:
The screen should be the shortest part of the experience.
RENDER APP
https://trailtutor-ai.onrender.com/
Local host:http://127.0.0.1:8000/
https://github.com/rajab-rajab/TrailTutor-AI
The public application is a lightweight FastAPI service deployed on Render.
For the live Gemma path, TrailTutor uses:
google/diffusiongemma-26b-a4b-it
through NVIDIA's hosted API.
A typical request contains:
{
"age": 13,
"environment": "school playground",
"topic": "plants",
"duration_minutes": 10
}
The response follows a small structured format:
{
"mission": "...",
"observe": "...",
"questions": ["...", "..."],
"safety": "...",
"reflection": "..."
}
This structure is intentionally restrictive. TrailTutor is not supposed to become another long conversation. Its job is to produce a useful mission quickly and then get out of the learner's way.
I wanted the core AI component to use an open-weight model rather than treat the model as a completely closed black box.
Gemma gave me a strong foundation for generating short, structured educational activities while keeping the architecture flexible enough to swap or self-host models later.
The live TrailTutor application uses DiffusionGemma for mission generation.
I also wanted to test whether a general open-weight model could be specialized specifically for TrailTutor's mission format.
I used Tinker to LoRA fine-tune:
Qwen/Qwen3.5-4B
The training setup was deliberately small and reproducible:
The dataset covers:
The held-out prompts have no exact input overlap with the training examples.
I evaluated the untuned and tuned versions of the same Qwen3.5-4B model on the 20 held-out prompts.
The final corrected evaluation produced:
| Metric | Untuned | Tinker-tuned |
|---|---|---|
| Overall structured-output score | 95% | 100% |
| Valid JSON | 19/20 | 20/20 |
| All required fields | 19/20 | 20/20 |
| Exactly two questions | 19/20 | 20/20 |
| Outdoor action present | 19/20 | 20/20 |
| Safety present | 19/20 | 20/20 |
| Reflection present | 19/20 | 20/20 |
| Mission under 70 words | 19/20 | 20/20 |
| Average latency | 3.355 s | 2.907 s |
That is a 5 percentage-point improvement in the held-out structural score.
Average latency also decreased by about 13.4%.
My first evaluation reported a larger improvement: 60% to 80%.
That result turned out to be misleading.
Some model outputs contained two consecutive valid JSON objects. My original parser took everything from the first opening brace to the last closing brace, then tried to decode the whole string as one JSON object.
That caused valid generations to be scored as failures with errors such as:
Extra data
Instead of hiding that mistake, I kept the original result in the repository and added a corrected v1.1 evaluation that parses the first complete JSON object while preserving the raw output.
The corrected result is the one I report here:
95% → 100%
I think keeping both versions is important because reproducible AI evaluation also means documenting when the evaluator itself was wrong.
TrailTutor is deployed publicly on Render using the project's Dockerfile.
Render hosts:
The computationally heavy Gemma inference happens through the model provider, so the web service itself stays lightweight.
This made it straightforward to turn the local prototype into a publicly accessible project that judges and users can try immediately.
User
|
v
TrailTutor web interface
|
v
FastAPI application on Render
|
+----------------------------+
| |
v v
DiffusionGemma Tinker experiment
via NVIDIA API Qwen3.5-4B
|
+--> untuned baseline
|
+--> LoRA-tuned model
TrailTutor benefits from open models because I can do more than simply send text to an opaque endpoint.
I can:
The Tinker experiment is a good example.
Instead of claiming that fine-tuning helped, I could actually train an open-weight model, evaluate it against its own untuned baseline, inspect failures, discover a bug in my evaluator, and publish the corrected evidence.
That kind of experimentation is much easier when the model ecosystem is open enough to adapt.
The biggest surprise was not the training result.
It was the evaluation bug.
Several outputs I initially counted as complete failures were actually good TrailTutor missions repeated twice.
That reminded me that model evaluation is not only about testing the model. The evaluator also needs to be tested.
It changed the final result from an apparent 20-point improvement to a more defensible 5-point improvement.
I prefer the smaller number because I can explain exactly where it came from.
I would like to continue in three directions:
I am entering TrailTutor AI for:
It is also automatically eligible for the overall Touch Grass challenge.
Many AI applications are designed to increase engagement time.
TrailTutor deliberately tries to do the opposite.
It uses AI to give a learner one useful reason to close the screen, step outside, pay attention to the physical world, and come back with something they noticed for themselves.