cd /news/machine-learning/failing-to-solve-manic-miner-with-rl · home › topics › machine-learning › article
[ARTICLE · art-140992] src=atomic14.com ↗ pub= topic=machine-learning verified=true sentiment=↓ negative

Failing to solve Manic Miner with RL

A developer's attempt to train a reinforcement-learning agent to complete the 1983 ZX Spectrum game Manic Miner failed, first with PPO over joystick inputs and then with move macros, according to a first-person account of the project. The developer and Claude pivoted to a two-pass search approach that first finds the optimum path and item order with guardians ignored, then searches along that route to wait and jump past guardians; the most expensive cavern, Return of the Alien Kong Beast, has five items and two switches, yielding 5,040 possible collection orders. The emulator runs about 1000 times faster than the original ZX Spectrum hardware.

read5 min views1 publishedSep 28, 2026

Around 10 years ago Google DeepMind showed that you could use deep reinforcement learning to play video games. It’s something that I found fascinating at the time and wanted to try but never had the time to really look into it.

With LLMs now being able to do a lot of the heavy lifting for me, I’ve started to work through my backlog of “this is interesting - I’d like to try it out”.

I grew up in the 70s and 80s and my childhood computer was the ZX Spectrum, and my ESP32 Rainbow project has given me a nice emulator that I can use to both play games and experiment with.

Manic Miner was one of the first games we purchased and I wasted a lot of time playing it. I have vague memories of progressing reasonably far, but never actually completing the game. So I naively thought - Reinforcement Learning - why not? Some cursory googling would have told me it was not suitable, but I’ve never really been one for doing a lot of up-front research before diving in.

The emulator that I have for the ESP32 also compiles into a desktop application (and there’s a WebAssembly version as well). Running it at full speed is about 1000 times faster than the original machine - so in theory we are in a good position to get something working.

There are also some really good resources for Manic Miner - this site is particularly good. Very handy when Claude is being an absolute doofus head.

The idea that Claude and I came up with (and of course Claude was super enthusiastic about…) was to wrap the emulator as a Gym environment, give the agent a reward for doing well, and run PPO until Willy walks out of Central Cavern. How hard can this possibly be…

This failed miserably. The initial attempt was to just use joystick movements to drive Willy around. We then tried creating a set of move “macros”, e.g. jump right, jump left, walk right. And still didn’t get anywhere with the RL approach.

Claude then suggested that what we were lacking was a good reward function, and that what we should do was solve the problem via search and then feed the route into the RL algorithm and use that to help guide the algorithm.

This really should have set off some alarm bells in my head.

We spent a lot of time messing around with searching and then feeding the path into the RL algorithm and started to get some results - but ultimately the question ends up being - “if we can solve this with search, what is RL adding?”.

So, eventually, after a lot of wasted time, we pivoted to just doing that - why don’t we just try and find the fastest route through each cavern using search.

Should be simple, shouldn’t it?

There are a couple of things that make it tricky. There isn’t a simple graph that we can search (see later to learn I was wrong). The search space is very large, we have moving guardians, sometimes quite a few of them, this explodes the search space. Every position Willy can end up in gets multiplied by all the possible guardian positions.

The order you collect items in can have a big influence on the length of the path - and you need to collect all the keys - this turns into a bit of a travelling salesman problem. What order should we collect them in?

Some caverns have switches that change the layout - once again the search space gets quite big.

Our first working version was a two pass system. The first pass runs with the guardians set as harmless and ignored. This lets us find the optimum path and item order.

The second pass takes this route and searches along it but waits and jumps to avoid the guardians.

The most expensive cavern is Return of the Alien Kong Beast, which has five items and two switches, and so 5,040 possible orders of collecting items. That one cavern took 8 hours to search.

Some of the caverns are very complicated, there are a lot of guardians, so waiting and jumping at the right time becomes critical.

This did get us a complete run through of Manic Miner - every cavern solved. 39,091 frames, which is about 13 minutes of play.

The problem was it took around a day of computer time to find it, and 92% of that time was spent in the emulator. Every move the search tried had to be run through a full Z80 emulation, and even at 1000 times real speed it’s pretty slow.

So we took the emulator out of the search. Willy’s movement is pretty simple - he walks 2 pixels per tick, jumps follow a fixed table, falls speed up and kill you if they’re too long, and crumbling floors and conveyors do what you’d expect. Claude went through the disassembly and wrote a model of all of this, and of the guardians, so the search can step through the game without running the emulator.

Most of the guardians don’t react to Willy at all - where they are just depends on how long you’ve been in the cavern. Eugene, the Kong Beast are the exceptions. So we could drop the harmless guardians pass and do a single A* search with the guardians live. The inner loop is in C++ and does around 690,000 states a second.

Every route the search finds still gets played back through the emulator, and if Willy dies we throw it away - this was added a belt and braces check - just in case Claude had got its model of the game wrong.

All the caverns are searched in order, each one starting from the state the previous one finished in.

There were bugs in the model. Eugene could walk off the bottom of his shaft, so Willy was dying to a guardian the model didn’t know was there. We also found that pressing fire to start the game was making Willy jump at the start of the first cavern.

The full run is now 36,791 frames, about 2,300 fewer than the two pass version, and takes about two hours to find instead of a day. Fifteen of the twenty caverns are faster and the other five are slower by between 2 and 40 frames.

Moral of the story? Forget all this ML and AI nonsense and get back to studying basic algorithms…

── more in #machine-learning 4 stories · sorted by recency
── more on @manic miner 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/failing-to-solve-man…] indexed:0 read:5min 2026-09-28 · —