{"slug": "failing-to-solve-manic-miner-with-rl", "title": "Failing to solve Manic Miner with RL", "summary": "A developer's attempt to train a reinforcement-learning agent to complete the 1983 ZX Spectrum game Manic Miner failed, first with PPO over joystick inputs and then with move macros, according to a first-person account of the project. The developer and Claude pivoted to a two-pass search approach that first finds the optimum path and item order with guardians ignored, then searches along that route to wait and jump past guardians; the most expensive cavern, Return of the Alien Kong Beast, has five items and two switches, yielding 5,040 possible collection orders. The emulator runs about 1000 times faster than the original ZX Spectrum hardware.", "body_md": "Around 10 years ago Google DeepMind showed that you could use [deep reinforcement learning to play video games](https://www.bbc.co.uk/news/science-environment-31623427). It’s something that I found fascinating at the time and wanted to try but never had the time to really look into it.\n\nWith LLMs now being able to do a lot of the heavy lifting for me, I’ve started to work through my backlog of “this is interesting - I’d like to try it out”.\n\nI grew up in the 70s and 80s and my childhood computer was the ZX Spectrum, and my [ESP32 Rainbow project](https://www.esp32rainbow.com/) has given me a nice emulator that I can use to both play games and experiment with.\n\nManic Miner was one of the first games we purchased and I wasted a lot of time playing it. I have vague memories of progressing reasonably far, but never actually completing the game. So I naively thought - Reinforcement Learning - why not? Some [cursory googling](https://www.mrkwatkins.co.uk/teaching-an-ai-to-play-zx-spectrum-games/) would have told me it was not suitable, but I’ve never really been one for doing a lot of up-front research before diving in.\n\nThe emulator that I have for the ESP32 also compiles into a desktop application (and there’s a WebAssembly version as well). Running it at full speed is about 1000 times faster than the original machine - so in theory we are in a good position to get something working.\n\nThere are also some really good resources for Manic Miner - [this site](https://skoolkit.ca/disassemblies/manic_miner/) is particularly good. Very handy when Claude is being an absolute doofus head.\n\nThe idea that Claude and I came up with (and of course Claude was super enthusiastic about…) was to wrap the emulator as a [Gym](https://gymnasium.farama.org/index.html) environment, give the agent a reward for doing well, and run [PPO](https://en.wikipedia.org/wiki/Proximal_policy_optimization) until Willy walks out of Central Cavern. How hard can this possibly be…\n\nThis failed miserably. The initial attempt was to just use joystick movements to drive Willy around. We then tried creating a set of move “macros”, e.g. jump right, jump left, walk right. And still didn’t get anywhere with the RL approach.\n\nClaude then suggested that what we were lacking was a good reward function, and that what we should do was solve the problem via search and then feed the route into the RL algorithm and use that to help guide the algorithm.\n\nThis really should have set off some alarm bells in my head.\n\nWe spent a lot of time messing around with searching and then feeding the path into the RL algorithm and started to get some results - but ultimately the question ends up being - “if we can solve this with search, what is RL adding?”.\n\nSo, eventually, after a lot of wasted time, we pivoted to just doing that - why don’t we just try and find the fastest route through each cavern using search.\n\nShould be simple, shouldn’t it?\n\nThere are a couple of things that make it tricky. There isn’t a simple graph that we can search (see later to learn I was wrong). The search space is very large, we have moving guardians, sometimes quite a few of them, this explodes the search space. Every position Willy can end up in gets multiplied by all the possible guardian positions.\n\nThe order you collect items in can have a big influence on the length of the path - and you need to collect all the keys - this turns into a bit of a travelling salesman problem. What order should we collect them in?\n\nSome caverns have switches that change the layout - once again the search space gets quite big.\n\nOur first working version was a two pass system. The first pass runs with the guardians set as harmless and ignored. This lets us find the optimum path and item order.\n\nThe second pass takes this route and searches along it but waits and jumps to avoid the guardians.\n\nThe most expensive cavern is Return of the Alien Kong Beast, which has five items and two switches, and so 5,040 possible orders of collecting items. That one cavern took 8 hours to search.\n\nSome of the caverns are very complicated, there are a lot of guardians, so waiting and jumping at the right time becomes critical.\n\nThis did get us a complete run through of Manic Miner - every cavern solved. 39,091 frames, which is about 13 minutes of play.\n\nThe problem was it took around a day of computer time to find it, and 92% of that time was spent in the emulator. Every move the search tried had to be run through a full Z80 emulation, and even at 1000 times real speed it’s pretty slow.\n\nSo we took the emulator out of the search. Willy’s movement is pretty simple - he walks 2 pixels per tick, jumps follow a fixed table, falls speed up and kill you if they’re too long, and crumbling floors and conveyors do what you’d expect. Claude went through the disassembly and wrote a model of all of this, and of the guardians, so the search can step through the game without running the emulator.\n\nMost of the guardians don’t react to Willy at all - where they are just depends on how long you’ve been in the cavern. Eugene, the Kong Beast are the exceptions. So we could drop the harmless guardians pass and do a single A* search with the guardians live. The inner loop is in C++ and does around 690,000 states a second.\n\nEvery route the search finds still gets played back through the emulator, and if Willy dies we throw it away - this was added a belt and braces check - just in case Claude had got its model of the game wrong.\n\nAll the caverns are searched in order, each one starting from the state the previous one finished in.\n\nThere were bugs in the model. Eugene could walk off the bottom of his shaft, so Willy was dying to a guardian the model didn’t know was there. We also found that pressing fire to start the game was making Willy jump at the start of the first cavern.\n\nThe full run is now 36,791 frames, about 2,300 fewer than the two pass version, and takes about two hours to find instead of a day. Fifteen of the twenty caverns are faster and the other five are slower by between 2 and 40 frames.\n\nMoral of the story? Forget all this ML and AI nonsense and get back to studying basic algorithms…", "url": "https://wpnews.pro/news/failing-to-solve-manic-miner-with-rl", "canonical_source": "https://www.atomic14.com/2026/09/28/manic-miner-from-reinforcement-learning-to-search", "published_at": "2026-09-28 00:00:00+00:00", "updated_at": "2026-09-28 13:18:27.881739+00:00", "lang": "en", "topics": ["machine-learning", "artificial-intelligence"], "entities": ["Manic Miner", "ZX Spectrum", "ESP32 Rainbow", "Google DeepMind", "Claude", "PPO", "Gymnasium", "Return of the Alien Kong Beast"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/failing-to-solve-manic-miner-with-rl", "markdown": "https://wpnews.pro/news/failing-to-solve-manic-miner-with-rl.md", "text": "https://wpnews.pro/news/failing-to-solve-manic-miner-with-rl.txt", "jsonld": "https://wpnews.pro/news/failing-to-solve-manic-miner-with-rl.jsonld"}}