cd /news/ai-agents/dharohar-an-open-source-ai-audio-gui… · home › topics › ai-agents › article
[ARTICLE · art-145955] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Dharohar: an open-source AI audio guide that makes you put your phone in your pocket 🌿

A developer in Indore, India built Dharohar, an open-source, audio-first heritage walking guide that uses a Mastra agent with SerpApi and Google's Gemma models to generate spoken scripts for nearby monuments and then deliberately blanks the screen so users walk without looking at their phones. The pipeline calls its tools in a fixed order because small Gemma models handle native tool calling unreliably, and ten npm evals enforce rules such as scripts staying under 100 words and never mentioning screens, apps or maps. The project runs Gemma 3 4B locally via Ollama or gemma-4-26b-a4b-it on Google AI Studio, with ElevenLabs eleven_multilingual_v2 voicing the guides.

by read5 min views1 publishedOct 6, 2026

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass I live in Indore. Ten minutes from my house is Rajwada, a seven-storey palace the Holkars started building in 1747. I have walked past it hundreds of times, usually looking at my phone.

Dharohar (धरोहर, heritage) is an audio guide for India's old city lanes that tries to get you to stop doing that. You tap Walk where I am. It finds the heritage sites around you, writes a short spoken guide for each one, voices it, and then asks you to put your phone in your pocket. The screen goes black. You walk. When GPS says you have reached the next stop, the guide starts talking.

The measure of success is how little you look at it.

It's for anyone in a heritage town who wants the story without the screen: residents who never learned the history of the street they live on, and visitors who don't want to follow a blue dot for an hour.

What keeps you off the screen:

Live: https://dharohar-agent.onrender.com. Open it on your phone, allow location, tap Walk where I am. (It runs on Render's free tier: the first visit after a quiet spell takes about a minute to wake up, and a brand-new stop takes about a minute to write.)

Video: https://youtu.be/wnoLYbbGi38?si=egkM_3cZ_uh14CGK In the video I'm standing at Holkar Stadium in Indore, a cricket ground with no heritage of its own. Dharohar looks wider, finds Gajanan Maharaj Temple, Indore Museum, Rajwada, Lalbagh Palace and Kanch Mandir, and orders them into a walk. Then: countdown, black screen, the guide begins. When I reach Rajwada, the phone vibrates and Rajwada's story plays.

You can also type any place. Gwalior Fort gives Chaturbhuj Temple, Gujari Mahal, Sasbahu Temple and the Siddhachal caves, all within 800 metres. Mandu gives Jama Masjid, Lohani Caves, Hindola Mahal and Nilkanth temple.

Audio-first heritage walking tours for Indian cities. A Mastra agent fetches live facts with SerpApi and Gemma writes a short spoken script with physical, real-world cues only. No screen talk: put the phone in your pocket and walk.

POST /script {"monument": "Rajwada, Indore"} → {"script": "...", "audioUrl": "/audio/rajwada-indore_<ts>.mp3", "duration": 32} GET /audio/<file>.mp3 serves the spoken guide (ElevenLabs eleven_multilingual_v2, so Hindi works too). audioUrl: null and a warning, and the script is still returned. gemma-4-26b-a4b-it) when GOOGLE_GENERATIVE_AI_API_KEY is set, else Groq when GROQ_API_KEY is set, else local Ollama (gemma3:4b). npm test, ElevenLabs mocked, no credits used) check audio, fallback and caching, and fail if a script runs over 100 words, mentions screens/apps/maps, or skips the… Each stop goes through a small pipeline, held together by a Mastra agent with two tools:

Then the walk page takes over: ordering stops into a route, geofences, wake lock and the offline cache.

Gemma, two ways. On my laptop Gemma 3 4B runs locally through Ollama, with no API at all. In production the same code calls gemma-4-26b-a4b-it on Google AI Studio. One environment variable switches between them, and Groq works too. A Gemma 4 script reads noticeably better than the 4B one. Kanch Mandir's script, for example, names Seth Hukumchand, gives the year it was built (1903) and points out the glass murals.

Why not let the model call the tools? Small Gemma models don't do native tool calling reliably. So the pipeline calls the tools in a fixed order instead of hoping the model remembers to search. That's simpler, and it never skips the facts.

Evals as the rules. The "touch grass" rules are tests, not hopes. npm test runs 10 evals:

GitHub Actions runs the fast evals on every push.

The eval that lied to me. My first version passed every check while Gemma described "a carved wooden balcony" and "a stone lion" at Rajwada. Neither exists. It had copied the example sentence from its own instructions. Format checks can't catch invented facts. What fixed it was grounding every script in search results, plus a rule to mention only features named in those facts.

Caching. Gemma on the free tier takes about a minute per new stop (the search itself takes 1.4 seconds). So each stop's script is cached, and its audio is cached per voice. The second person to walk past Rajwada gets it in under a second, and costs no model call and no ElevenLabs credits.

Agent tracing. The pipeline is instrumented for Sentry: each tour is one agent run, with separate spans for the search, the Gemma call (model, tokens, latency) and the voice. I found the one-minute Gemma step by timing each part from outside; with tracing switched on, that breakdown shows up per request.

I built this with an AI coding agent (Claude Code). I set the idea, the rules and the evals. The agent wrote most of the code, ran the tests and fixed what failed, and I reviewed every change. The session is below.

Local history belongs to locals. A temple trust, a heritage walk group or a college club can run Dharohar on its own hardware. I ran the whole thing on my laptop with Gemma through Ollama, with no API key for the model at all.

Swapping models is one line. This happened during the build: the Gemma version I first deployed was retired on one provider. Because the model is open-weight and the code doesn't depend on one vendor, I listed what was available and switched to Gemma 4 in one line. With a closed API, that's a rewrite or a shutdown.

Open data, open code. Walks come from Wikipedia's open geodata. The guide is grounded in public search results. The code is MIT-licensed, so anyone can add a walk for their own town.

What's next: a "local story" box at each stop, so people who live there can add the oral history that never reached Wikipedia, read aloud as part of the walk. That only works if the community can see and change how the guide works. That's what open is for.

render.yaml blueprint.

── more in #ai-agents 4 stories · sorted by recency
── more on @dharohar 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/dharohar-an-open-sou…] indexed:0 read:5min 2026-10-06 · —