Show HN: Bench-bench – Long-horizon planning, measured in bench press Bench-bench v0.2, a public benchmark created by an independent developer, measures AI models' performance on long-horizon planning tasks, humorously framed as 'bench press' capability. The project aims to test whether AI models can handle sustained, multi-step reasoning, addressing a gap in current evaluation methods. Bench-bench v0.2 · public run How much can an AI model bench press? I keep hearing labs talk about training, lift, and moving weights, but can their models even lift ? Bench-bench is my attempt to test that. bro