cd /news/ai-agents/claude-code-mods-minesweeper-and-tes… · home › topics › ai-agents › article
[ARTICLE · art-144539] src=dev.to ↗ pub= topic=ai-agents verified=true sentiment=↑ positive

Claude Code mods: Minesweeper and testkit

A developer built Minefield, a Minesweeper mod for Claude Code that opens a game pane beside the transcript while the agent works, shipping with 31 tests run via `claude plugin test` on two pinned Claude Code builds. The developer found that four bugs passed the test kit but only surfaced in real sessions, including a pane that could not take keyboard focus because Claude Code 2.1.285 refuses a `focus: true` request while the command's own text is still in the prompt. The fix re-requests focus after `/mines` returns, retrying at 250 ms, 500 ms and 1000 ms while checking `$.ui.panes()` so a pane the user closed is never reopened.

by read16 min views2 publishedOct 3, 2026

My own harness, the tooling I run around Claude Code, comes with a custom trust score. That trust score decides whether I can leave the agent alone yet, without supervision or if it's in my best interest to babysit it. As you see in the header image, it's babysitting era, so I built a Minesweeper to kill time between turns.

From Claude Code 2.1.287, mods load by default. A mod is a plugin whose hooks are JavaScript or TypeScript functions running inside Claude Code instead of shell commands, and it can open and draw its own panes. Minefield is Minesweeper as a mod: /mines opens a pane beside the transcript, and you clear the board while Claude works.

It is one 9×9 board with 10 mines, a game for the few minutes Claude works. The first reveal is always safe and opens an area, clicking a fully flagged number opens the cells around it, and the best time is kept between sessions. The orange face follows your cursor. Stop for six seconds and it pretends to read the transcript beside it, eyes half shut, line by line. It only looks; the mod sees nothing of the transcript. It winks at flags and puts on sunglasses when you win. This was the most fun part to implement.

It ships with 31 tests that run through claude plugin test, and CI runs them on two pinned Claude Code builds. Four of its bugs passed the test kit anyway and only showed up in a real session: the pane couldn't take the keyboard, the hotkeys did nothing, the face wrapped, and the mouse wasn't there.

In claude plugin test, nearly every engine call your mod makes ( ui.open, ui.panes, store.get) is answered by an on(...) stub you write in the test file; the kit answers a few itself, such as $.ui.invalidate. That is why it runs fast and without credentials. It also means the host in your tests behaves exactly as well as your doubles do.

/mines opened the pane, and the keys kept going to the prompt.

Seen in live sessions on 2.1.285: a pane takes the keyboard only over an empty prompt. Anthropic's docs list focus: true from a command as one way a pane gets the keyboard, but while command.run runs, the command's own text is still in the prompt, and on 2.1.285 the request was refused. The pane opens and the keyboard stays where it was.

In the tests, ui.open is answered by my own double, and my double says yes:

  on('ui.open', (_$, e) => {
    opened.push({ rows: e.rows, columns: e.columns, closeOnEscape: e.closeOnEscape })
    return { value: { isPlaced: true } }
  })

The fix is to open the pane, then ask for the keyboard again once /mines has returned. register.ts waits 250 ms, then 500, then 1000 between tries (FOCUS_RETRIES_MS), three tries inside two seconds, and every retry first asks $.ui.panes() whether the pane is still open and still without the keyboard. If the pane is gone, the person closed it, and a retry must not bring it back.

// Ask for the keyboard again once /mines has returned, while the pane is open
// and still without it; never reopen a pane the person has closed.
function focusSoon($: EngineInterface, attempt: number): void {
  const wait = FOCUS_RETRIES_MS[attempt]
  if (wait === undefined) return
  try {
    $.clock.after(wait, () => void refocus($, attempt))
  } catch {
    // No timer; a click on the pane gives it the keyboard.
  }
}

async function refocus($: EngineInterface, attempt: number): Promise<void> {
  try {
    const pane = (await $.ui.panes()).find((p) => p.id === PANE)
    if (pane === undefined || pane.isFocused) return
    await openPane($)
    focusSoon($, attempt + 1)
  } catch {
    // The panes cannot be listed; a click on the pane gives it the keyboard.
  }
}

The retry logic is testable, with mock.clock standing in for $.clock:

test('a pane the person closed is not reopened by the focus retries', async ($, on) => {
  const opened: Opened[] = []
  const clock = world(on, { opened, pane: { isOpen: false, isFocused: false } })
  await $.session.start(SESSION)
  await $.command.run(mines())
  await clock.advance(5000)
  expect(opened.length).toBe(1)
})

What no test here can check is the refusal itself. isFocused in that test is whatever my double says it is.

A pane has two places to draw from. The hooks module (register.ts) answers the ui.render hook and returns elements: boxes, text, buttons. A Client is a surface module (board.ts) with its own state, a frame timer, and pointer and key handlers, which is where a game wants to live. So the control buttons went in the board, next to the cells they act on, each with a one-letter hotkey.

None of the hotkeys fired. A Button's hotkey fires only for buttons the hooks module draws, never for buttons a Client draws.

The tests had no way to notice. ui.press({ key: 'flag' }) presses a button by its element key ('flag', where the hotkey is f), so a press in a test never goes through the hotkey at all.

The fix moved the buttons into register.ts. That raised the next problem: the buttons now live in the hooks module, the game lives in the board, and the board learns about presses only through its props. Several presses can arrive between two redraws of the board, so handing over only the latest move drops the others. Each press gets a number instead, and the hooks module keeps the last 32:

// Moves kept for the board, so presses faster than its redraws all arrive.
const MAX_ACTS = 32
js
function ask(type: MoveType): void {
  const n = (state.acts.at(-1)?.n ?? 0) + 1
  state.acts = [...state.acts, { n, type }].slice(-MAX_ACTS)
}

The board plays every move newer than the last one it played, in order, once:

// Every move newer than the last one played, in order. Presses can outrun
// the board's redraws, so one redraw may bring several. Numbers lower than
// the last one played mean the hooks module started counting again, as it
// does when the plugin reloads, so every move it lists is new. The same game
// back when there is nothing new.
export function applyActs(game: Game, acts: readonly Act[]): Game {
  const last = acts.at(-1)?.n ?? 0
  let next = last < game.acted ? { ...game, acted: 0 } : game
  for (const move of acts) {
    if (move.n > next.acted) next = { ...act(next, move.type), acted: move.n }
  }
  return next
}

A mounted test can't catch the race this guards against. The kit redraws after every press, so it never coalesces presses, and a mounted test of the queue passes even with the bug in place. So the race is tested as a pure function, with the presses that land in one redraw written out by hand:

test('every move newer than the last one played is played, in order, once', async () => {
  const ready = newGame(SEED)
  // Three presses that arrive in one redraw, the first one already played
  const acts = [
    { n: 1, type: 'right' },
    { n: 2, type: 'right' },
    { n: 3, type: 'down' },
    { n: 4, type: 'reveal' },
  ] as const
  const played = applyActs({ ...ready, acted: 1 }, acts)
  expect(played.cursor).toEqual({ x: 5, y: 5 })
  expect(played.acted).toBe(4)
  expect(played.marks[5 * 9 + 5]).toBe(OPEN)
  // Nothing newer: the same game back
  expect(applyActs(played, acts)).toBe(played)
  // After a reload the hooks module counts its moves from 1 again: every move it lists is new
  const reloaded = applyActs({ ...played, acted: 15 }, [{ n: 1, type: 'left' }])
  expect([reloaded.cursor, reloaded.acted]).toEqual([{ x: 4, y: 5 }, 1])
})

The face is four rows of text, nine columns wide (FACE_WIDTH = 9), drawn in Claude Code's own orange: the theme's claude colour, so it follows your theme. The status sits beside it in a row of boxes.

A row of boxes shrinks every box to fit. In a real pane the face's column came out one column short, and each of its nine-column rows wrapped in two. Eight rows of orange, and a face that looked like something had sat on it.

The tests find the face by its text (eyes(' ● ● '), a ui.find for that Text in the board), and the text was right. The strings were nine columns wide; the column holding them was eight.

The fix is a fixed width plus flexShrink: 0 on the face's column. The status column beside it shrinks and truncates (wrap: 'truncate-end'), so the long status is the part that gets cut, and the face's own rows truncate rather than wrap:

  // The face's column keeps its width, so long text beside it is cut and the
  // face is not.
  const face = faceRows(game).map((row) => Text({ ...row.style, wrap: 'truncate-end', children: [row.text] }))
  const faceColumn = Box({ flexDirection: 'column', width: FACE_WIDTH, flexShrink: 0, children: face })

The board's own rows carry flexShrink: 0 for the same reason.

On Claude Code's main screen, the pane gets no clicks. Claude Code asks the terminal for pointer reporting only in the fullscreen layout: CLAUDE_CODE_NO_FLICKER=1 at start, or /tui fullscreen in a session. On the main screen it asks for none at all.

In the kit, ui.pointer({ type: 'down', ... }) hands the board a click directly. In a real terminal the click has to be asked for first. To see whether Claude Code asked, tmux display-message -p '#{mouse_any_flag}' prints the terminal's mouse flag.

So the game is keyboard-first, and every action has a key. The letter keys (n new, r reveal, f flag, w a s d to move) are the hotkeys of the control buttons from the previous bug, drawn by the hooks module, so they work as soon as the pane has the keyboard, with no click. Arrow keys, h j k l, space and return reach the board through onKey, which a Client gets only after a click gives it the keyboard. Escape hands the keys back to the prompt.

The line under the buttons says which screen you are on:

// One line under the controls, for the screen the pane is on: the mouse
// exists only in the fullscreen layout, and a pane without the keyboard
// takes them back with ctrl+x tab.
function hintFor(isFocused: boolean, isFullscreen: boolean): string {
  if (!isFocused) return 'ctrl+x tab gives the pane the keys again.'
  if (isFullscreen) return 'Click reveals, right-click flags; click a fully flagged number to open around it.'
  return 'Keys only on this screen: /tui fullscreen adds the mouse.'
}

So the test kit gets the same treatment as the agent: I don't leave it alone either. Before a mod counts as done, it runs in a real interactive session, driven by a script. tmux does the driving: start Claude Code detached at a fixed size, type the command, send mouse escapes for clicks, and read the screen back as text.

tmux new-session -d -s mod -x 160 -y 50 "CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 CLAUDE_CODE_NO_FLICKER=1 claude --plugin-dir <mod>"
tmux send-keys -t mod -l '/<command>'; tmux send-keys -t mod Enter
tmux send-keys -t mod -l $'\e[<0;COL;ROWM'; tmux send-keys -t mod -l $'\e[<0;COL;ROWm'   # left click (SGR); button 2 is right
tmux capture-pane -t mod -p        # the screen as text; -e keeps the colours

A few notes on it:

CLAUDE_CODE_ENABLE_FUNCTION_HOOKS=1 is the switch mods need on 2.1.285 and 2.1.286. From 2.1.287 they load by default.\e[<0;COL;ROWM is an SGR mouse press at that column and row, and the lowercase m is its release. Swap the 0 for a 2 to right-click.capture-pane -p prints the screen as plain text, which means Claude can read its own UI back and check it. One rule for the loop: never script a win. The best time lives in the plugin's $.store, which survives sessions, so a scripted win writes a record to the real store, and my best time would belong to tmux. The live sessions played through a loss. The wins, and the stored best time, are covered by tests with a store double.

claude -p "/plugin-types" writes .claude/types/claude-code.d.ts for the installed build, so one pinned Claude Code build serves the typings, the engine and the test kit, and the three can't disagree. Typecheck against those, and leave the typings on the main branch of anthropics/claude-code alone: the two drift. On main, ui.open answers void. On 2.1.285 it answers { isPlaced }.

The typings stay out of git. Claude Code ships under "All rights reserved" (its npm LICENSE.md), and a copy in an MIT repo would breach that:

.claude/types/
*/.claude-plugin/types/
*/tsconfig.json

/plugin-types runs without a login on 2.1.285, behind the switch. From 2.1.287, claude -p asks for a login. So CI runs two jobs on every push to main and every pull request. The first writes the typings, typechecks, validates and tests on 2.1.285:

  typecheck:
    runs-on: ubuntu-latest
    env:
      CLAUDE_CODE_VERSION: 2.1.285
      CLAUDE_CODE_ENABLE_FUNCTION_HOOKS: '1'
      DISABLE_AUTOUPDATER: '1'
    steps:
      - uses: actions/checkout@v7

      - uses: actions/setup-node@v7
        with:
          node-version: 22

      - name: Install Claude Code
        run: npm install -g "@anthropic-ai/claude-code@${CLAUDE_CODE_VERSION}"

      - name: Write the mod typings for that build
        run: claude -p "/plugin-types" < /dev/null

      - name: Typecheck
        run: npx -y -p typescript@5.9.3 tsc -p .

      - name: Validate and test each game
        run: |
          for manifest in */.claude-plugin/plugin.json; do
            game="${manifest%/.claude-plugin/plugin.json}"
            claude plugin validate --strict "$game"
            claude plugin test "$game"
          done

The second installs 2.1.288, where mods are on by default, validates the marketplace, and runs claude plugin validate --strict and claude plugin test again, on the kind of build players run.

Then there's what the mod can see. The host reads on(...) and $.noun.method(...) from the hooks module's source before it runs any of it, which is why register.ts spells them literally and keeps every helper that takes $ at the top level. claude plugin validate ./minefield prints what it read:

❯ ./register.ts hooks: session.start, command.run{command=mines}, ui.render{component=Pane}, ui.message
❯ ./register.ts calls: $.clock.after (via focusSoon), $.clock.now (via deal), $.command.register, $.session.surfaces, $.store.get, $.store.set (via recordWin), $.ui.invalidate, $.ui.open (via openPane), $.ui.panes (via refocus), $.ui.resolve
❯ ./register.ts surface modules: hooks/board.ts

Four hooks: session start, its own /mines command, pane drawing (it draws its own and passes every other pane on untouched), and messages from its board. No prompt, tool or attachment hook, and no network call. The "By Reporails" at the end of the status line is a Link element; your terminal opens it, and the mod makes no request.

When I pulled the awesome-claude-code-mods data on 3 October 2026, 479 of its 1,018 mods were listed as seeing every prompt or every tool call. Plenty need to: the tool-call counter on Anthropic's own mods overview has to see every tool call. A Minesweeper has no use for either. The count, if you want to re-run it:

curl -sL https://raw.githubusercontent.com/karanb192/awesome-claude-code-mods/main/data/mods.json | python3 -c "import sys,json;m=[x for x in json.load(sys.stdin)['mods'] if x['kind']=='mod'];print(sum(any(k in ' '.join(x['sees']) for k in ('every prompt','every tool call')) for x in m),len(m))"

In a wide fullscreen layout the pane docks beside the transcript and opens floor to ceiling, at the 60 columns it asks for. A 9×9 board at one row a cell uses a corner of that. fitFor picks the biggest cell the room holds, from the pane's columns and rows, biggest cells first and the face under the board before beside it:

// Every fit, biggest cells first; under the board before beside it.
const FITS: readonly Fit[] = [
  ...SCALES.flatMap((k) => [
    { width: 2 * k, height: k, side: false },
    { width: 2 * k, height: k, side: true },
  ]),
  { width: 3, height: 1, side: false },
  { width: 3, height: 1, side: true },
  { width: 2, height: 1, side: false },
  { width: 2, height: 1, side: true },
  { width: 1, height: 1, side: false },
  { width: 1, height: 1, side: true },
]

// The biggest fit the room holds; with the room unknown, or too small for
// any, a row a cell, two columns wide where the room has them, the face under
// the board.
export function fitFor(cols: number, rows: number, room: { columns: number; rows: number }): Fit {
  const fit = FITS.find((f) => columnsFor(f, cols) <= room.columns && rowsFor(f, rows) <= room.rows)
  return fit ?? { width: cellWidthFor(cols, room.columns), height: 1, side: false }
}

SCALES is [3, 2], so the big fits are cells six columns by three rows, or four by two. A terminal cell is about twice as tall as it is wide (the reason a board at one row a cell already spends two or three columns a cell where it can), so 2k columns by k rows is square on screen. The gaps need the same care. Each tile is width - 1 columns of colour with a one-column gap after it, and its last row is an upper half block, ▀, so the gap between rows is half a row tall, the same size as the one-column gap between tiles. At the six-by-three fit that's five columns by two and a half rows of colour, square again:

  const across = fit.width - 1
  const glyphRow = Math.floor((fit.height - 1) / 2)
  const glyphAt = Math.floor(across / 2)
  for (let y = 0; y < game.rows; y++) {
    const tiles: Tile[] = []
    for (let x = 0; x < game.cols; x++) tiles.push(tileOf(game, x, y))
    for (let r = 0; r < fit.height; r++) {
      const runs: Run[] = []
      for (const { glyph, ink, fill } of tiles) {
        if (r === fit.height - 1) pushRun(runs, '▀'.repeat(across), { color: fill })
        else if (r === glyphRow) pushRun(runs, ' '.repeat(glyphAt) + glyph + ' '.repeat(across - glyphAt - 1), { ...ink, backgroundColor: fill })
        else pushRun(runs, ' '.repeat(across), { backgroundColor: fill })
        pushRun(runs, ' ', {})
      }
      out.push(runs)
    }
  }

Since fitFor is a pure function, the sizes are tested without a pane:

  // A docked pane: square tiles three rows tall, or two, where they fit
  expect(fitFor(9, 9, { columns: 60, rows: 36 })).toEqual({ width: 6, height: 3, side: false })
  expect(fitFor(9, 9, { columns: 60, rows: 25 })).toEqual({ width: 4, height: 2, side: false })

Above the prompt, rows are what's scarce, so the pane asks for the rows of the most compact fit the terminal's width holds: one row a cell, a hidden tile drawn as a short box (▆▆), and the face and the status beside the board where they fit there.

In a Claude Code session:

/plugin marketplace add reporails/arcade
/plugin install minefield@reporails-arcade

Then /mines. It needs Claude Code 2.1.287 or later. For the docked pane and the mouse, start Claude Code with CLAUDE_CODE_NO_FLICKER=1, or run /tui fullscreen. If minefield doesn't load, run claude plugin test in an empty folder: no hooks module to load means mods load for you. hooks modules are turned off here means a setting blocks them: disableAllHooks in your settings, or your organization's policy. hooks modules are turned off in this process means Anthropic has them off for your account, and no local setting changes that.

The repo is MIT, reporails/arcade on GitHub. Post your best time in the comments.

── more in #ai-agents 4 stories · sorted by recency
── more on @claude code 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/claude-code-mods-min…] indexed:0 read:16min 2026-10-03 · —