# Claude and Jev Play Pokémon

> Source: <https://claude-and-jev-play-pokemon.singlepageapp.co/>
> Published: 2026-09-28 21:34:10+00:00

## What we learned

The short version: Jev made good choices when it was told the truth, Claude was a better detective than planner, and most of the thirty hours went on problems in how the bot saw and moved through the game, not in how it chose. Over the run, the line between the two models also moved. By the end, Claude's code was making most of the decisions and Jev was mostly agreeing with it. The rest of the page has the details.

### Jev

#### Good at

- Choosing sensibly from correct facts. Asked what to do next with the team under half HP, it chose to heal 56 times out of 56.
- Reading type matchups. In 3,506 battle turns it never picked a move that couldn't hurt the enemy when one that could was on offer.
- Using evidence. Shown the runs where paralysis had lasted all the way to the Champion, it bought a Full Heal.
- Being consistent. Asked the exact same question twice, it gave the same answer 99.4% of the time.

#### Weak at

- Weighing the goal against the facts. The goal said to reach Giovanni on 11F, and in the Silph Co. lift it chose 11F in 11 of 19 questions, even with "DEAD END… visited 3x with no progress" written next to it.
- Knowing what things are worth. Asked to clear space in the bag, it threw away all three Rare Candies rather than any of nine spare TMs.
- Letting go. About 165 times it picked an option already marked as failed when an untried one was on the list.
- Long lists. With 16 or more options, over half the questions ended in a near-tie.

None of the long stalls were Jev's fault. In the worst one, at Vermilion, it was asked the same question 1,471 times, and every option on the list was marked as failed.

### Claude

#### Good at

- Finding the ground truth. When the map made no sense, the deep sessions went to the game's ROM and the pokered disassembly. That's how the teleport pads, arrow tiles and locked doors were solved.
- Measuring instead of guessing. At the Elite Four it ran the same code from several save states and counted the wins: 3 of 5.
- Knowing when not to ship. The session whose code won changed nothing, and another threw away a change that won 1 test in 8.

#### Weak at

- Fixing the right layer. Before every big stall was fixed, the short sessions changed what Jev was shown, adding facts and hiding options, while the real bug was in how the bot moved.
- Doing the job it kept putting off. 36 session summaries said the bench needed training. None did it, and the Elite Four came down to a coin flip.
- Keeping its notes true. `NOTES.md` has at least eight corrections marked "WRONG", and some coordinates were written "from memory, not from the game data".
- Tidying up. 4,587 lines of bot code were added and 747 removed, and 175 lines name a specific map.

Bit by bit, Claude took over the deciding. In the winning run, when a where-next question had one option labelled "THE WAY FORWARD" or "BEST", Jev picked it 1,672 times out of 1,699. When Jev kept disagreeing, the next session removed the option it preferred. The improver was told that Jev makes every strategic decision, but nothing stopped the facts turning into advice. Christian's project has a commit for exactly this: "Facts, not advice".

### The harness

#### Worked

- Escalating. After three sessions without progress, the next one got thirty minutes and was told that "a position that stays the same while a direction is pressed means the bot's picture of the map is wrong… not that Jev chose badly". Seven of the eight badges came after one.
- The test gate. A session had to test its change and get an answer from Jev, or the change was thrown away. It blocked 11 sessions, and all but one deserved it.
- Save states. Claude could test from the exact spot the bot was stuck, and measure win rates.
- Showing what the game knows. Exposing Mt Moon's elevation walls in `grid()` got the bot through the cave; an earlier attempt without them spent about an hour and a half there.

#### Didn't

- Measuring progress. The story-flag count includes one flag per trainer beaten, so the bot looped in Silph Co. for an hour while the count crept up and no deep session started. Team strength wasn't counted at all.
- The map. `grid()` was wrong in four ways that each stalled the bot: arrow tiles shown as plain floor (the host decoded them but never passed them on), only one kind of tile counted as water, the old map returned for a moment after every warp, and switch-opened walls shown as closed after a battle. Claude had to find and patch each one, and its water patch later sent the bot back and forth on Route 23.
- Keeping records. A bug replaced two deep sessions' summaries with "summary above stands", and that is what the next eight sessions read.
- Memory. Every reload, about every five minutes, wiped whatever the bot had worked out while running.

### What we'd take to the next project

1. **Get the agent's view of the world right first.**`grid()` was the bot's only view of the game, and it was the biggest single thing holding the bot back. It showed arrow tiles as plain floor, counted only one kind of tile as water, returned the old map for a moment after every warp, and showed walls opened by a switch as still closed after a battle. Claude spent deep session after deep session working out that the map was wrong and patching around it, and one of those patches caused a new loop. When an agent loops, check what it can see before blaming how it chooses, and if the environment already knows the truth, hand it over.
2. **Guard the line between facts and advice.** If one model writes the facts and another decides, the writer ends up deciding unless something stops it.
3. **Measure the progress you actually want.** Counting trainer flags hid an hour-long loop, and not counting team strength left the Elite Four to a coin flip.
4. **Make the long sessions automatic.** One-minute fixes covered most of the miles but never got past a real block.
5. **Work you keep putting off gets more expensive.** "Train the bench" sat in the notes for twenty hours.

## The whole run

The game tracks 261 story flags, starting with picking your starter and ending at the Hall of Fame. This line counts how many were set at the end of each five-minute run. The red ticks are badges and the shaded bits are where the bot got stuck. The wobble at the end is the Elite Four: every time the bot lost there, its flags reset.

## Eight badges

Seven of the eight badges needed a deep session first. That's a longer Claude session that has to work out what's wrong before it writes any code. They all found much the same thing: Jev's choices were fine, but the bot had the map wrong. Each picture is the last frame of the run where that badge was won.

1. 
### Brock Pewter City, 1 h 36 mThe first time it got stuck, it went back and forth between Viridian City and the Poké Mart. Three sessions in a row blamed Jev. Then the first deep session found three real bugs: the walk-to-edge code stopped one row short of the map seam, the south exit got marked as failed after eight ticks even if the bot had moved, and the map data lagged a moment behind every warp. Brock went down to a Lv15 Bulbasaur.
2. 
### Misty Cerulean City, 3 h 32 mMt Moon took two hours and four blackouts. The bot's money fell from ¥1,929 to ¥120 in twenty minutes, the least it ever had. Every ladder on the bottom floor goes up to the floor above, and the way the bot merged doors left Jev with exactly one option: the ladder it had just come down. Then a Zubat finished off Ivysaur at 11% HP, because the heal option had told Jev a wild fight only costs 8%.
3. 
### Lt. Surge Vermilion City, 5 h 42 mOnce it had Cut, the bot spent fifty minutes wandering around Route 11 and Diglett's Cave. In one run it asked Jev the same where-next question 1,522 times. When it finally cut down the tree in front of the gym, it gave up on the door instead of walking in, and the tree grew back.
4. 
### Erika Celadon City, 7 h 47 mRock Tunnel got the bot an Onix and a Geodude, its first catches since the starter, and a twenty-minute freeze. In trainer battles it answered yes to "Will RED change POKéMON?" and then picked the Venusaur that was already out. Again and again.
5. 
### Koga Fuchsia City, 13 h 0 mThe longest stretch, and it got stuck three different ways. The arrow tiles in the Rocket Hideout look like plain floor to the bot, so it slid across the same one for half an hour. At the Route 12 gate it pressed the Poké Flute's menu buttons before the menu had opened, and one run went nineteen minutes without asking Jev anything. Then on Route 14 it always crossed at the one row that drops you into a four-tile pocket behind a Bird Keeper.
6. 
### Sabrina Saffron City, 17 h 44 mSilph Co took 2 h 43 m and five deep sessions. You can't walk to the Card Key's corridor at all; the only way in is a teleport pad two floors up. Locked doors open when you face them and press A, not when you walk into them. Sabrina's room is walled in on every side, and you reach it through a chain of pads. The bot could only walk, until one session pulled every pad's landing square out of the ROM.
7. 
### Blaine Cinnabar Island, 22 h 15 mLapras knew Surf and the bot had the badge to use it, but as far as the bot knew, water was a wall. Once that was fixed it surfed down into Seafoam Islands, went round in circles on the lower floors, then sat in a sealed-off patch of sea west of the islands, pressing left into a rock. In the end it gave up on the sea and went to Cinnabar the long way round, back through Lavender, Pewter and Pallet Town.
8. 
### Giovanni Viridian City, 22 h 30 mThe quickest badge: fifteen minutes, three sessions, $1.04. The session just before it spotted that Jev was still being told to go and get badge 7 from Blaine, and wrote new goals for badges 7 and 8.

## Twelve tries at the Elite Four

It's five fights back to back with no healing in between, and every blackout halves your money. Each column is one attempt, with Lorelei at the bottom and the Champion at the top. Filled rooms were beaten. The red room is where that attempt ended.

| Room | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 | 11 | 12 | 
|---|---|---|---|---|---|---|---|---|---|---|---|---|

The League took seventeen deep sessions, and each one found something new. The bag was full, so the shop sold nothing and the bot turned up at Agatha with no potions. The two Haunters had the same level and HP, so the bot read the wrong one's PP. Solarbeam and Mega Drain got used up in the wrong rooms. Max Potions went on the Haunters, while Venusaur, revived to half HP, never got offered one. The shop only bought Max Potions, which can't revive anything. Nothing could cure paralysis. And the damage estimates ignored type, so a Jynx hit estimated at 28 actually did 112.

The code that finally won went in at 00:52. The three sessions after it tested it, found it won about half the time, and left it alone.

The live pass is a coin flip.

Session summary, 01:27 UTC, seven minutes before the win

## Who decided what

### Jev picked

Whenever the bot handed over a choice, it gathered up whatever options it had at that moment, attached a few facts to each, and asked Jev. How much it handed over shrank as the run went on (see [What we learned](#h-lessons)). Up to the win it asked 17,673 times, always in the same format. Each answer took about a third of a second and cost $0.000037, so the whole lot came to $0.65.

| Where should the player go next on this map? | 10,292 | 
|---|---|
| Which move this turn? | 3,780 | 
| What next: progress, heal, train or shop? | 2,263 | 
| Wild battle: fight, switch or run? | 593 | 
| Keep the current Pokémon or switch? | 497 | 
| Everything else: elevators, shops, TMs, boulders, the starter | 248 | 

About a fifth of Jev's answers were near-certain, and in about one question in nine the top two options were almost tied. Even the starter was close: Bulbasaur 50%, Squirtle 47%.

Here's one of those questions in full. The bot asked it on Route 22, just after the rival's Charizard had beaten it three times in a row. The fight is described at the top, Venusaur's four moves are the options, and Jev's pick is at the bottom.

That line about Venusaur having no damaging move is wrong. Cut still had PP and did normal damage to Charizard, and it had the best score of the four. The bot's check for damaging moves had a bug, so Jev was told there weren't any, and with that in front of it, Jev went for Leech Seed. A later session caught it and fixed the check.

### Claude fixed

There were 304 sessions. 257 of them took about a minute and cost around $0.31 each, doing small jobs: relabelling the exits for the next leg, adding a fact, fixing a menu. If three sessions in a row made no progress on badges, story flags, Pokémon owned or maps visited, the next one became a deep session. That meant thirty minutes, finding the cause before writing any code, and proving the fix with a test that got past the block. There were 47 deep sessions. They were 15% of the sessions but 54% of the $176.52 spent.

Every long stall turned out to be a bug in how the bot moved or saw the game. Deep sessions were told not to trust the diagnoses in the notes without evidence, and nearly every one started by ruling Jev out.

| What was actually wrong | For example | 
|---|---|
| The bot's picture of the map was incomplete | Arrow tiles in the Rocket Hideout. Teleport pads to the Card Key and to Sabrina. Locked doors need A, not walking. A rock read as water while surfing. | 
| Stale reads and caches | Map data from before the warp. A cached destination re-run without pressing anything. Failures remembered per map name on a map split in two by a gate. | 
| Menu timing | The Poké Flute pressed before its menu opened. "Will RED change POKéMON?" answered yes forever. The lobby clerk not responding after a loss. | 
| Missing mechanics | Surfing. Pushing boulders. Healing between Elite Four rooms. Curing a status problem. Throwing something away when the bag is full. | 
| Wrong facts given to Jev | "The boulder switch stays pressed." "A wild fight costs 8% HP." Damage estimates with no type bonus. Goal text still naming a badge already won. | 
| Resource strategy | Which room gets Solarbeam. Revive against Max Potion. A Lv86 carry with a Lv22 to Lv40 bench. | 

- The block was a set of bugs in how the bot moves, not bad choices by Jev. 01:02 UTC, Vermilion
- Arrow tiles look like plain floor in grid(), and grid() has no arrow data, so the bot's pathfinder kept routing to the B4F stairs over (18,16). 04:07, Rocket Hideout
- Sabrina stands in a room with walls on every side. The only way in is a chain of teleport pads, and the bot could only walk. 13:15, Saffron Gym
- Pushing the 1F boulder is undone once the player leaves 1F, and the heal option told Jev the opposite. 21:15, Victory Road
- I found no physical block, so I shipped no fix: the loop is Elite Four and Champion fights that the current code wins only about half the time. 00:52, the session whose code won

## What the bot became

The host runs one file, `bots/bot.js`, in a V8 isolate: no Node, no network and a 128 MB memory cap. It gives the bot a handful of synchronous calls and reloads the file whenever it changes. That's all Claude had to work with.

```
// Passed in, and also available as globals. Every call blocks until it is done.
export default function play({
  press, hold, wait, frame,        // buttons, a frame at a time: A B START SELECT UP DOWN LEFT RIGHT
  state, party, box, bag, battle,  // read straight from the game's memory, never written
  screen, menu, sprites, event,    // text rows, dialog, the ▼ prompt, the cursor, story flags
  grid,                            // this map: walls, grass, water, ledges, doors, NPCs, blocked moves
  read, sym, rom,                  // bytes by pokered symbol; species, moves, items, maps, type chart
  jev,                             // jev(state, questions) → { answers, usage }, billed to the host
  save, load, snapshot, log,
}) { /* ... */ }
```

The bot sees every map as text, from `grid()`. Here's the legend it had to work from.

```
PALLET_TOWN 20x18
 ##########,,########      .  walkable            ,  tall grass
 #..........@.......#      #  wall or obstacle    ~  water (needs SURF)
 #...####....####...#      T  cuttable tree       C  counter (talk across it)
 #..##D##...##D##...#      v < > ^  ledge, jump over it that way
 #..N...............#      D  door, stairs, warp  N  NPC or object   @  you
```

### The starting point

A fresh start drops in a 40-line bot. It presses through Oak's intro, takes the preset names RED and BLUE, and stops in Red's bedroom. Everything else the bot does was added after this.

```
export default function play({ press, wait, screen, state, log, snapshot }) {
  if (state().playerName === 'RED') { log('already past the intro: nothing to do'); return; }
  for (let i = 0; i < 400; i++) {
    const s = screen(), st = state();
    if (st.playerName === 'RED' && st.map === 'REDS_HOUSE_2F' && s.nonMapTiles === 0) break;
    if (isNamePrompt(s)) { press('DOWN'); press('A', 4, 20); continue; }
    press(s.hasTextBox || s.cursor ? 'A' : 'START', 4, 20);
  }
  wait(60);
  snapshot('bedroom');
}
```

### The loop it ended with

Thirty hours later [the file](https://github.com/stephenheron/claude-and-jev-play-pokemon/blob/main/bots/bot.js) was 3,699 lines long, plus a 35-line helper, and it's still one function with one loop. Every tick it reads the state, deals with whatever's on screen, and if nothing's in the way, asks Jev where to go next and walks there. Here it is cut down to the bones. The names are real, except the menu line and the heal line, which each stand in for a couple of hundred lines of special cases.

```
export default function play() {
  if (gstate().playerName !== 'RED') intro();
  for (let tick = 0; tick < 100000; tick++) {
    const st = gstate(), s = screen(), p = party();
    if (moveLearnStep()) continue;               // "wants to learn X": Jev picks which move to forget
    if (st.inBattle) { battleStep(); continue; }  // pickMove, askSwitch, catch or RUN
    if (st.map !== loopMap) { settleMap(st.map); continue; }   // grid() lags a warp by a few frames
    if (openMenuWithNoQuestion(s)) { press('B'); continue; }   // the flute bug, never again
    if (s.hasTextBox || s.waitingForA || s.cursor) { press('A'); continue; }
    if (maybeLobbyShop(st) || maybeFieldHeal(st) || maybeTeachTM()) continue;
    const gl = goal(st, p);                       // Jev: progress, heal, train or shop
    if (gl === 'heal') { walkToNearestCenter(st); continue; }
    if (solveVRBoulder(st) || elevatorStep(st)) continue;
    goDest(st);                                   // Jev picks a door or a map edge; code walks there
  }
}
```

Underneath that loop are the pieces the deep sessions added, most of them a few dozen lines: a breadth-first pathfinder over `grid()` that treats ledges and the Rocket Hideout arrow tiles as jumps, a table of Saffron Gym teleport pads read from the ROM, a surfing mode that makes water walkable, a boulder-pushing planner for Victory Road, a damage estimate for every move Jev gets offered, and a PP budget for the five League rooms. There are 204 comments in the file that start with a session id, one for each fix, and 12 of them are marked as the root cause.

### How it asks Jev

The helper file has three functions. This is the one the bot uses whenever it asks Jev something, and it's called from 21 places. The context is a plain object, the options map names to facts, and the fallback is what happens if Jev is switched off or won't answer.

```
export function askChoice(context, instructions, options, fallback) {
  try {
    const { answers } = jev(context, { pick: { type: 'choice', instructions, criteria: options } });
    const a = answers.pick;
    log(`jev: ${a.choice} (conf ${a.confidence})`);
    return a.choice in options ? a.choice : fallback;
  } catch (e) { log(`jev failed: ${e}`); return fallback; }
}
```

Here's what those 21 places ask, in the order they appear in the file: the Cinnabar Gym quiz, the starter, which move this turn, whether to switch after a faint, whether to switch when a trainer sends out something new, what to do next, which item to throw away, what to buy in the Indigo lobby, whether to use a healing item before the next Elite Four room, which floor in the Silph Co elevator, which floor in the Rocket Hideout lift, where to go next on this map, which move to forget (two places), whether to swap a boxed Pokémon in, which TM to teach, which member gets it, which boulder goes on which switch (two places), whether Lapras learns Surf, and who learns Strength.

### How it grew

Lines of bot code at each milestone. Almost all of it came from deep sessions, and hardly any of it was ever deleted.

| First run | 40 | 
|---|---|
| Badge 1, Brock | 467 | 
| Badge 2, Misty | 649 | 
| Badge 3, Lt. Surge | 992 | 
| Badge 4, Erika | 1,166 | 
| Badge 5, Koga | 1,678 | 
| Badge 6, Sabrina | 2,291 | 
| Badge 7, Blaine | 2,812 | 
| Badge 8, Giovanni | 2,833 | 
| Entering the Elite Four | 3,303 | 
| The winning run | 3,734 | 

The sessions also kept a running log in `bots/NOTES.md`, a section each, which reached 3,334 lines. Every session had to edit a copy of the bot, run it against the real game from the live save, and get at least one answer from Jev during that test. Otherwise its changes were thrown away, which happened eleven times.

## Every blackout halves your money

The bot won ¥432,138 in prize money and ended up with ¥32,234. Every drop is a loss. The big ones are Mt Moon, losing to the rival on Route 22 three times in half an hour, and the Elite Four.

## Trainer card

The bot mashed A through the nickname screen for Bulbasaur and for everything it caught after that. So the player is RED, the rival is BLUE, and all six Pokémon are called AAAAAAAAAA. The game said "RED got on AAAAAAAAAA!" 94 times, once for every surf.

- TIME
- 29:53
- MONEY
- ¥32234
- BADGES
- 8
- POKéDEX
- 10 OWN 119 SEEN
- A PRESSES
- 532316
- BATTLES
- 1161
- RAN AWAY
- 504
- ZUBAT MET
- 176
- CUT USED
- 618
- THUNDERBOLT
- 3

- AAAAAAAAAAVENUSAUR L89
- AAAAAAAAAAGEODUDE L22
- AAAAAAAAAAHAUNTER L40
- AAAAAAAAAAHAUNTER L38
- AAAAAAAAAAHAUNTER L33
- AAAAAAAAAALAPRAS L18

Nobody stopped the bot after the credits, so it just kept going. It walked back to the League, beat the Champion again at 02:57, and then spent five hours in Victory Road, where Venusaur hit Lv99.
