This is my submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.
TimeWalk is a location-aware historical storytelling backend. You give it a GPS coordinate and a question like "What is the history of this place?" β it gives you back a short, source-grounded story about where you're standing.
It's built for travelers, walkers, and anyone who has ever stood in front of a building and wondered what happened there before. Open the app, get your location, get the story.
The tagline I kept coming back to while building it:
Where am I, what happened here, and can the AI prove it?
The "Touch Grass" framing fits exactly. The point isn't to keep you on your phone β it's to give you just enough context in just enough time that you look up, look around, and actually see the place you're in. Sections are short. Language is plain. Sources are there when you want them, tucked at the bottom. Nearby places are things you can walk to.
It's a backend-only submission for now. The mobile frontend is the next milestone.
In the video, I run a live request for Kasba Peth, Pune (18.5204Β°N, 73.8567Β°E). Watch for:
whyHere line for each, and sources at the bottom
Every section in the output is grounded in a source the model actually retrieved. Nothing is fabricated.
Repo: github.com/tanish/timewalk (https://github.com/tanishbhongade/timewalk)
Structure:
src/
βββ agent/ # Two-phase LLM agent (research + synthesis)
βββ controllers/ # HTTP edge
βββ middleware/ # Auth, rate limiting, validation, errors
βββ providers/
β βββ geocoder/ # Nominatim adapter
β βββ search/ # Tavily adapter, Wikipedia-first
β βββ llm/ # Bedrock factory
βββ routes/
βββ schemas/ # Zod schemas for everything
βββ services/ # Domain logic
Stack: Node.js + TypeScript + Express, AWS Bedrock for the LLM, Tavily for web research, Nominatim for reverse geocoding, Zod for validation, JWT for auth.
I'm using google.gemma-3-27b-it on AWS Bedrock. Open weights, 27B parameters.
The honest comparison from my testing:
The key architectural decision, and it took me a while to get right:
Phase 1 β Research. The LLM gets three tools and a bounded budget (AGENT_MAX_TOOL_CALLS=6). It decides what to search for and calls the tools. No output format constraints. Its only job is to gather evidence.
Phase 2 β Synthesis. A second LLM call with no tools. Its only job is to turn the gathered evidence into a strict JSON response.
This is where the "can the AI prove it?" promise gets kept. Three enforcement mechanisms:
This is a hard boundary. It makes fabrication structurally impossible, not just discouraged.
Tavily's includeDomains parameter lets me restrict the first search to Wikipedia. If Wikipedia has a page for the place, that's the only result the model sees β one clean tier-1 source. If it doesn't, the search falls through to a broader query.
The result: clean sources for well-covered places, and graceful degradation for the long tail.
Three bugs worth mentioning, because each one taught me something:
Bedrock rejected my synthesis prompt with The toolConfig field must be defined when using toolUse and toolResult content blocks. I'd been replaying the raw tool-call conversation into the synthesis call, and Bedrock's API is stricter than OpenAI's about orphaned tool blocks. Fixed by extracting evidence into a plain text block instead of replaying messages.
The model returned sources objects without id fields, breaking schema validation. Fixed by making id optional in the schema and deriving it deterministically from the URL when missing.
Confidence labels drifted to "high" for everything, including claims the model's own caveats said were tradition-based. Fixed by adding a prompt rule: if the narration hedges, the confidence must not be "high."
The real reason open weights mattered for this project: I could actually compare models on the same pipeline and pick the one that worked. I didn't have to commit to a vendor before I knew whether the model was good enough. I ran DeepSeek, Ministral, and Gemma through identical prompts, identical tools, identical responses β and Gemma 3 27B won on the thing that actually mattered (grounding).
With a closed API, that comparison would have cost me three integrations and several days of glue code. With open weights served through an OpenAI-compatible interface, it was a one-line env var change.
Three other reasons this build leans on open:
Cost predictability. A hobby project can't run up a $500 OpenAI bill in a weekend of testing. Open weights on a pay-per-token host let me iterate without worrying about the meter.
Data privacy. A location-aware app is a privacy-sensitive product. Users are telling me where they are, in real time. With open weights, I have the option to self-host later β the queries never have to leave infrastructure I control. That's an option I'd lose with a closed model, and it's one I want to keep open as the app grows.
Iteration speed. I changed the prompt maybe 40 times during development. Every one of those iterations had to fire a model call. Being able to swap models, adjust temperature, and try different providers without re-architecting anything is what made the whole two-phase design discoverable at all.
The "can the AI prove it?" framing only works if you can test the AI. Open weights gave me that.
Best Use of Gemma
One last thing. The reason this project exists isn't that I wanted to build a history app. It's that I wanted to build something that makes a person standing on a street corner look up and see a place differently. Every architectural decision β the short sections, the whyHere line, the sources tucked at the bottom β exists to serve that moment.
If you build something like this, I'd love to hear what you learn. Especially if you try it in a place without a Wikipedia page. That's the real test.