cd /news/ai-agents/touchgrass-a-phone-referee-for-outdo… · home › topics › ai-agents › article
[ARTICLE · art-149303] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

TouchGrass: a phone referee for outdoor games, with an open-weight Game Master

A developer built TouchGrass, an open-source phone-based referee for outdoor games such as Tag and Hide & Seek, pairing a deterministic game engine with an AI Game Master powered by the open-weight Gemma model. The Game Master only proposes twists like safe zones or bonus challenges, and each proposal must pass schema validation, a game legality check, and deterministic safety validation before the engine applies it, so play continues if the model is slow, malformed, or disabled. Tag and Hide & Seek are playable for 2–4 players, while the browser gameplay client remains unimplemented and the project is MIT licensed.

by read7 min views1 publishedOct 11, 2026

This is a submission for the Hacktoberfest Open-Source AI Challenge Week 1: Touch Grass.

TouchGrass is a referee for outdoor games that lives on your phones, so nobody has to stay behind to keep score.

The idea is simple: everyone should be able to play instead of arguing about the rules, keeping score, or watching the clock. TouchGrass is designed to handle those jobs so friends, families, and youth groups can spend more time playing outside.

The project has a game engine, a multiplayer server, persistent game sessions, and an AI Game Master powered by an open weight model.

The Game Master can propose a twist for a round: a safe zone rule, a bonus challenge, or a little more time. But there is one important distinction: the AI proposes; the game engine decides.

The model never directly changes the game state. Its proposal must pass three checks implemented in ordinary code before it can be applied:

Gemma proposes
      ↓
Schema validation
      ↓
Game legality check
      ↓
Deterministic safety validation
      ↓
Game engine applies the event
      ↓
Players receive the update

If the model is slow, produces malformed output, proposes an illegal event, or is switched off, the game can continue without that AI event.

The project is intended for friends, families, and youth groups who want a reason to go to the park, and for anyone who has ever argued about whether a tag counted.

What works today: Tag and Hide & Seek are playable for 2–4 players, with three rounds and a named winner. Three other games Ice & Fire, Treasure Hunt, and Nature Challenge have their rules defined but are not playable yet. The website labels them as coming soon rather than pretending they are finished.

The browser-based gameplay client is not implemented yet. The game flow has been tested locally through the game server.

Live website: https://touchgrass-game-master.vercel.app/

The live website presents the project, games, rules, and design. It does not yet support playing a complete game directly in the browser.

Local demo: The game server has been tested locally with multiple players and a real Game Master turn. The recording below demonstrates session creation, player joining, a round starting, and a Game Master event being applied after validation.

GitHub repository: https://github.com/SHWETANK-SAURABH/touchgrass_game_master

The project is MIT licensed. The README contains the commands needed to install dependencies, run verification, start the server, and try the local demo.

The Game Master uses three important pieces:

In my local setup, Gemma runs on my own machine. No hosted AI API is required for a Game Master turn in that configuration.

One rule shaped the entire architecture:

The AI proposes. The safety layer validates. The game engine executes.

A language model can be wrong, slow, or creative in ways that are inappropriate for an outdoor game. It should not be responsible for deciding whether an action is legal or whether the game state is valid.

Instead, TouchGrass uses a deterministic game engine. Game events are checked against the rules, and the same valid event sequence produces the same game-state transitions. That makes the core mechanics testable independently of the model.

The model is treated as an untrusted source of suggestions, not as the authority over the game.

Each turn follows a controlled pipeline:

game_history retrieves historical game records, while game_rules provides information about the rules and constraints. Tool calls cannot directly mutate game state. Unknown tools are rejected, and tool inputs and outputs are bounded.

If the model produces malformed output or the proposal fails validation, the game does not apply the invalid event.

This design lets the AI contribute creativity without giving it control over the rules.

The project uses:

The repository has more than 2,000 automated tests, including 35 scripted Game Master evaluation scenarios. The goal is to make changes measurable rather than relying only on whether a prompt sounds better.

The production server currently uses a single process multiplayer architecture. It is not designed for multiple independent game server instances sharing the same database.

1. The model saw something it should not have.

During development, a game join code leaked into the model's view through a historical game record. The model did not need that information. I traced the issue to the history view and corrected the redaction.

That experience reinforced an important lesson: limiting the model's authority is not enough. The information it receives must also be limited to what it needs.

2. Small models can be cautious.

In my tests with Gemma 3 (4B), the Game Master frequently chose to extend the round. That is a legal and safe option, but repeated use makes the game less varied.

A smaller model can be useful for a controlled task, but its behaviour still needs evaluation. The next challenge is to encourage more variety without weakening the validation rules.

3. Safety rules have limits.

The safety validator is a set of deterministic rules. It can reject hazards that it recognises, but it does not know the park where players are standing, the traffic nearby, or the actual conditions around them.

It does not make outdoor play automatically safe. Players still need to use their judgement and follow local safety rules.

4. A successful model response is not the same as a successful game event.

The model can return valid JSON and still suggest something that is illegal in the current game. Structured output solves only one part of the problem.

Keeping schema validation, engine legality, and safety validation separate made those failure cases easier to test and reason about.

Three things became easier because I could run and inspect an open-weight model.

Because Gemma runs locally through Ollama in my development setup, I can inspect what goes into the model and what comes back.

That helped me discover the join-code leak and correct the model's input.

Open weights do not automatically guarantee privacy or security, but they give developers more control over the inference environment and make it easier to investigate model behaviour.

The Game Master evaluation suite contains 35 scripted scenarios covering expected outcomes and failure stages.

Running these tests against a fixed model version makes comparisons more meaningful when I change prompts, tool descriptions, or safety rules. It does not eliminate model variability, but it gives me a repeatable baseline.

In my local configuration, the Game Master runs through Ollama without making a per-turn call to a hosted AI provider.

That can make experimentation more accessible to developers, families, and youth groups who do not want a usage-based AI bill.

It is important to distinguish local inference from a completely offline application: dependencies and model weights must be installed first, and other services may still need network access.

The broader benefit is freedom to experiment. Someone else can add a game, replace the model, improve the evaluation suite, or translate the safety rules without needing permission from a closed API provider.

Gemma 3 (4B) acts as the Game Master. Through Ollama, it receives game context, can consult read-only tools, and proposes a single game twist as structured JSON.

The architecture is built around what a small open-weight model can contribute ideas while leaving legality and execution to deterministic code.

The Game Master agent is built using Mastra. Its two read-only tools, game_history and game_rules, provide additional context before it proposes an event.

The application does not give the agent unrestricted control over the game. Its output must pass the existing schema, engine-legality, and safety checks before any change is applied.

The immediate priorities are to build the browser-based gameplay client, make the remaining three games playable, and improve the variety of Game Master events without compromising the safety boundary.

I also want to make the project easier for open-source contributors to extend through better game definitions, more evaluation scenarios, accessibility improvements, and localisation.

I built TouchGrass with help from an AI coding assistant, including assistance with implementation and this article. The architectural decisions, testing, and decisions about what the project can honestly claim remain my responsibility.

The central idea is simple: use open-weight AI to make games more interesting, but never let the model become the referee.

Let the model suggest. Let code check. Let everyone play.

── more in #ai-agents 4 stories · sorted by recency
── more on @touchgrass 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/touchgrass-a-phone-r…] indexed:0 read:7min 2026-10-11 · —