cd /news/ai-agents/workshop-brief-battleship-with-human… · home › topics › ai-agents › article
[ARTICLE · art-125481] src=gist.github.com ↗ pub= topic=ai-agents verified=true sentiment=· neutral

Workshop Brief: Battleship with Human-Steered AI

A workshop brief outlines an experiment in which a human ensemble steers an AI coding agent to build a Battleship game engine with a separate text client, rather than letting the AI work autonomously. The brief requires the AI to draft plans, designs, code, and tests while humans retain responsibility for rules, API, architecture, and test decisions, and to surface open questions in a clarification session instead of assuming standard Battleship rules. It also mandates a shared Spec Loop DDD glossary and defers bonus features like variable board size and two-player mode until the main part is accepted.

by read7 min views20 publishedAug 28, 2026

This file contains the original workshop specification used during the workshop. See the workshop analysis and the revised version below.

This brief is shared unchanged with workshop participants and the AI coding agent.

This workshop is meant for a human ensemble working together, not for a single human occasionally approving the AI's work.

The human ensemble may add its own instructions first, or ask the AI to start with a clarification session.

This workshop is an experiment in using an AI coding agent without turning the human ensemble into passive observers.

The human ensemble stays responsible for the whole result. The humans do not need to type everything, but they do need to steer the work, review it, and accept it together.

The AI may draft plans, designs, code, and tests, but it must not quietly take over important decisions about the rules, the API, the architecture, the design, or the tests.

Build a Battleship game engine that is separate from any presentation layer.

The required client is a small text client used to test and play the engine.

After the main part is accepted, the group may optionally add a browser demo on top of the same engine. The browser demo is out of scope until the group explicitly brings it into scope.

  • Start with a one-player game.
  • Use a 10x10 board.
  • Use this fleet, described only by length and count:
    • 1 ship of length 4
    • 2 ships of length 3
    • 3 ships of length 2
    • 4 ships of length 1
  • Do not rely on standard ship names from another version of Battleship.
  • Any ship may be marked as a submarine.
  • Submarines may occupy different depths.
  • Only depth charges can hit submarines.
  • The engine and the text client must stay separate.

This workshop is supposed to include real discussion. Not everything is fixed up front.

Important parts of the rules, the engine API, the architecture, the class-level design, and the test design are intentionally left open. The group should discuss them. The AI should actively surface them in a clarification session.

If an unclear point could change behavior, the public API, the design, the tests, or the task breakdown, the AI must ask instead of silently assuming a standard Battleship rule. The AI should not waste time asking about obvious, trivial, or already settled points.

This workshop must use a shared glossary.

Use it as a Spec Loop DDD glossary: a shared domain language for the workshop, not a dump of implementation terms.

Use the glossary for important game-rule, API, and design terms that need one stable meaning across clarification, task files, design, and test specification.

Reuse existing agreed terms instead of inventing local synonyms. When a clarification session settles a term or changes its meaning, record it before later work depends on it.

Do not let the same word quietly mean different things in different parts of the workshop. Do not put implementation-local details into the glossary unless they are part of the reviewed contract.

The main part is the smallest reviewable version that satisfies the workshop baseline. It includes:

  • a game engine that is separate from the client
  • a small text client that can play against that engine
  • starting a game
  • showing enough game state to play
  • performing attacks
  • enough support for submarine behavior that the engine/client boundary still makes sense

The exact engine API is not fixed here. The engine must support starting play and performing attacks, but the concrete public contract is part of the workshop design.

Keep these as bonus features only. Do not plan or implement them before the main part is accepted.

  • variable board size

  • game statistics

  • two-player mode

  • showing statistics on demand

- two-player terminal play
- player-specific prompts
  • showing both boards at the end of each turn in two-player mode

The human ensemble is responsible for the whole result: the rules, the API, the architecture, the design, the implementation direction, the test intent, and the final acceptance.

The AI may write artifacts and code, but that does not move responsibility away from the human ensemble.

This is the default flow. The group may adapt it, but the review gates below must still be respected.

  1. The human ensemble discusses the game, the workshop goal, and any extra instructions.
  2. The AI runs a clarification session for the current scope.
  3. The human ensemble confirms the clarified baseline for the current scope.
  4. The AI creates a vertical breakdown as separate task files.
  5. The human ensemble reviews that breakdown.
  6. The AI prepares the current task file in implementation-ready detail.
  7. The human ensemble reviews and approves the current task's scope, class-level design, and exact test specification.
  8. The AI implements and tests the current task.
  9. The human ensemble reviews the result.
  10. Repeat for the next task.

Because this workshop uses separate task files and no subtasks, the human ensemble may choose the working mode separately for each task.

For a given task, one AI session handles the work in order.

If an important change appears during implementation, it goes back to planning and human review before implementation continues.

For a given task, the group may work on later tasks while an implementation agent works on the current approved task.

In parallel mode:

  • each reviewable slice gets its own task file
  • do not use subtasks for this workshop
  • the planning agent may edit task files, but not code
  • the implementation agent may edit code and add Implementation notes , but must not otherwise edit task files
  • no worktrees are assumed
  • if the implementation agent makes an important deviation to keep moving, it must record that deviation in Implementation notes and bring it back for explicit human accept/reject review before the task is accepted
  • after humans accept a deviation, the group must make sure the task files and other planning artifacts are brought back into sync before later work depends on them

This workshop uses Spec Loop explicitly.

The AI must:

  • begin with a clarification session by using spec-loop-clarify-task
  • use a clarification session before planning whenever important open decisions remain
  • actively surface important open decisions about rules, API, architecture, design, and test specification
  • after the clarification session, stop for human confirmation of the clarified baseline before creating task files
  • avoid fake safety theater: do not present artificial alternatives for obvious points
  • use the task-file path, not chat-only planning
  • use the Spec Loop DDD glossary rules and keep the shared domain language updated for important terms
  • use separate task files for separate reviewable vertical slices
  • avoid subtasks in this workshop
  • split work into small reviewable slices of behavior whenever a credible split exists
  • avoid planning-only, test-only, or layer-only implementation tasks unless the humans explicitly ask for that structure
  • prepare only the current task in full implementation-ready detail
  • before asking to implement the current task, make sure humans have reviewed and approved:
    • the current task scope
    • the class-level design
    • the exact test specification
  • use spec-loop-prepare-execution-approval before asking to implement the current task
  • use spec-loop-implementation-flow after implementation is approved
  • in sequential mode, send important implementation-time changes back to planning and human review before continuing
  • for each task, allow the human ensemble to choose sequential or parallel mode
  • in parallel mode, a workshop-specific override is allowed: the implementation agent may keep moving through some important deviations, but it must record them in Implementation notes and bring them back for explicit human review before the task is accepted

The implementation stack is intentionally not fixed in advance.

The group may choose it during the session. If the stack choice could change the design, the tests, or the task shape, the AI should raise it in the clarification session rather than assume it.

A browser demo may be added only after the main part is accepted. Until then, do not treat it as part of the main scope.

By the end of the workshop, the group should have:

  • a clarified baseline for the current scope
  • a shared glossary for the current scope
  • a task-file breakdown into small reviewable slices
  • an implementation-ready current task with class-level design and exact test specification before coding starts
  • reviewed code and tests for completed tasks
  • a working engine plus text client before any bonus work begins

Do not start by coding.

Start with a clarification session for the current scope. Resolve what is already settled. Ask about important open decisions. Do not silently fill important gaps with assumed standard Battleship rules. After the clarification session, stop for human confirmation of the clarified baseline. Then create or update task files, prepare the current task for human review, and wait for approval before implementing.

── more in #ai-agents 4 stories · sorted by recency
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/workshop-brief-battl…] indexed:0 read:7min 2026-08-28 · —