Flawed benchmarks (epoch.ai benchmark registry filtered by flawed) Epoch AI maintains a registry of 85 AI benchmarks spanning mathematics, software engineering, agentic workflows, and games, and has flagged a subset as flawed. The registry includes benchmarks such as a long-horizon coding test in which models reimplement entire programs without access to the original source code, a 100-puzzle chess set generated programmatically with a chess engine, and 1,000 factoid questions covering politics, science and technology, art, sports, geography, and music. The registry also aggregates many benchmarks into a single general capability scale and lists collections of unsolved research mathematics problems, including problems posed by mathematician Paul Erdős that remained open as of August 2026. A registry of 85 AI benchmarks covering mathematics, software engineering, agentic workflows, games, and more. An index aggregating many different benchmarks into a single, general capability scale. A collection of unsolved problems from research mathematics, curated to be especially interesting and difficult. A collection of problems posed or studied by mathematician Paul Erdős, open as of August 2026 and curated to be especially interesting and difficult. A long-horizon coding benchmark. Models reimplement entire programs end-to-end, without access to the original source code. A test of AI systems' learning capabilities, measuring whether their scores improve across repeated playthroughs of the relatively obscure game Earthborne Rangers. 100 positions from a puzzle-oriented variant of a well-known game, the details of which are deliberately left undisclosed. 100 novel puzzles, generated programmatically with a chess engine. Each puzzle has a single best next move. 1,000 factoid questions about politics, science and technology, art, sports, geography, music, and more. Expert-written math problems covering advanced undergrad through early career research.