This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass
Most trail apps ask you to look at them. At a junction you take the phone out, wake the screen and find the blue dot, and for that moment you are looking at a map instead of the trail. This week's theme asks for builds where the screen is the shortest part of the experience, so I built a trail guide that you listen to.
Waymark turns a real trail into a short spoken guide. You pick a trail, or open a GPX file, and Waymark works out the facts from open data: where the junctions are and which way to go at each one, where the climbs start and how steep they are, where the water, bridges, steps and viewpoints are, and how long the walk should take. Gemma 4 E2B, an open-weight model running inside the browser, writes one or two plain sentences for each of those waymarks and a short briefing for the whole walk.
Before you leave, a 3D model of the real terrain previews the walk in about a minute, with every cue shown where it will be spoken. Then you send the guide to your phone and put the phone in your pocket. Pocket mode keeps the screen dark, follows GPS and speaks each cue when you reach that waymark, for example: "Continue straight onto Mist Trail. John Muir Trail on the right is not your route." On a two-hour walk, you look at the screen for about a minute.
It is for walkers and hikers on trails they do not know well, and for anyone who would rather keep their eyes on the path. A group can share one link.
Live app: https://starknightt.github.io/waymark/
Things to try:
The phone view during a simulated walk. Each cue is also shown as text when it is spoken:
Gemma writing a guide for the Triund trail in the browser, with nothing mocked (sped up 2 times):
Each cue next to the facts it came from:
A trail guide you listen to. Pick a real trail. Waymark builds the facts from OpenStreetMap and open elevation data, and Gemma 4 E2B, running in your browser on WebGPU, writes a short spoken cue for every junction, climb and landmark. A 3D model of the real terrain previews the walk in about a minute. Then the phone goes in your pocket: it keeps the screen dark and speaks each cue when you reach that waymark.
Live: https://starknightt.github.io/waymark/
Built for the DEV Hacktoberfest Open-Source AI Challenge, Week 1 (Touch Grass). The repository was started on October 5, 2026, inside the challenge window.
TypeScript, Vite, Three.js and LiteRT-LM, under the MIT license. All of the code was written during the challenge window, starting on October 5.
A spoken guide is worse than useless if it says left when the path goes right, and a 2B model cannot do geometry. So the work is split. Code builds the facts:
These are merged into waymarks, between 5 and 22 on the twelve test trails, each a short list of facts. The start, the finish and every junction where the route turns or moves onto a different path are marked as required. This is what Gemma sees for one waymark on the Mist Trail, wrapped here to fit:
w8 (REQUIRED): junction: continue straight onto "Mist Trail";
"John Muir Trail" on the right is not your route
gemma-4-E2B-it-web.litertlm (2.0 GB). The browser downloads it once into its private file storage (OPFS), and after that it loads in a few seconds (2.3 s on my PC).write_guide({ cues: [{ waymark, text }], briefing }). The waymark field is an enum of the real waymark IDs, so the model cannot invent a waymark, and constrained decoding guarantees the call can be parsed.
Every cue is checked against the facts of its own waymark before it is saved:
A sentence that fails is removed. If nothing safe is left, the cue is rebuilt from the facts. The checker has 35 unit cases that run in Chrome, and most of them come from real mistakes.
add_cue once per waymark and set_briefing once. It called write_guide call with an array fixed it.THREE.Color converts colours to linear light, so my tree placement compared sRGB pixels with linear values and matched nothing. Every mountain was bare until I read the hex value directly.
Twelve trails in seven countries, Chrome on an RTX 4060 desktop GPU, greedy decoding. The second column asks for the same guide as JSON in plain text, without constrained decoding.
| One constrained tool call (shipped) | JSON in plain text | |
|---|---|---|
| Cues written by Gemma | 180 | 180 |
| Passed every check unchanged | 171 | 170 |
| Changed by the checker | 9 (8 of them on Mount Takao) | 10 |
| Required waymarks covered | 120 of 120 | 120 of 120 |
| Output that could not be used | 0 | 0 |
| Time per guide | 18.2 s | 11.4 s |
| Decoding speed | 44 tokens/s | 80 tokens/s |
Eight of the nine changes were on Mount Takao, where Gemma translated or romanised Japanese path names, for example "Takao Mountain Line" for 高尾山線. The checker replaced them with the names on the map and kept the turns.
The result I did not expect: Gemma 4 E2B followed the JSON format in plain text on every trail, so constrained decoding mainly bought certainty, at a cost of about 45% of the decoding speed. I kept it, because a guide for a trail nobody has tested must always parse. Most of the accuracy came from the facts and the checker rather than from the decoding method.
Two more checks. On the live site, in a fresh browser, the first guide took about 10 minutes, nearly all of it the 2.0 GB download on my connection; writing the guide itself took about 20 seconds. All 62 cues of the four demo guides were also checked line by line against their source facts during development. After the checker, none of them gives a wrong turn, distance or climb, though a few read awkwardly, such as "There is a bridge, Mist Trail.", where the bridge is mapped as part of the Mist Trail. One mistake the checker cannot see is in a briefing: the Pen y Fan briefing says the walk reaches the summit of Pen y Fan, while the map data puts the summit 250 m from the route.
Where you walk, and when, is personal. In January 2018, Strava's public heatmap of its users' activity revealed the locations and patrol routes of military bases. A trail plan says where you will be and when you will be away from home. With an open model running in the browser, the route, the guide and your position never go to an AI provider.
It has to work without signal. Trails are often where coverage fails. The guide is written before you leave and fits in a link of about 1.5 KB, so on the trail the phone needs neither a network nor a model.
I could control how the model writes, not only what I ask it. Constrained decoding with my own schema, where the only valid waymarks are the real ones, and a 12-trail evaluation that I could re-run as often as I liked, at no cost. Most of the mistakes described above were found by those repeated runs. With a paid API I would have run the evaluation less often and found fewer of them.
It is open all the way down. OpenStreetMap data, open elevation tiles, an Apache-2.0 model, an Apache-2.0 runtime and MIT code. Anyone can host Waymark as static files, swap in another model, or adjust the facts for the trails where they live.
The trade-off is a 2.0 GB first download and a desktop GPU to write guides. A larger hosted model would write more varied sentences, but for spoken directions being right matters more than style, and that comes from the facts and the checker, which work the same way with an open model on my own machine.
Best Use of Gemma. Gemma 4 E2B is the only model in Waymark, and it runs entirely in the browser through LiteRT-LM on WebGPU. It writes every spoken cue and every briefing through a constrained tool call whose schema only allows real waymarks. Every demo guide on the site was written by it, and "Write it again on this device" runs it live.
Thanks for reading. If you try Waymark on a trail near you, I would like to hear how the cues held up.