Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires A blog post examines three categories of benchmarks — performance "napkin math" estimates, AI model evals including DeepSWE and Senior SWE-Bench, and winter tire comparisons — arguing each is misused as evidence. The post highlights the sirupsen/napkin-math GitHub repository (5.4k stars) and its latency and throughput tables, then asks what is wrong with each benchmark. The piece contends that people cite DeepSWE and Senior SWE-Bench to claim one AI model is better than another, and that winter tires are commonly asserted to beat all-season tires in cold weather. We're going to look at three different kinds of benchmarks, one set of calculations for baseline numbers for performance "napkin math" estimates, one set of AI model evals, and one on car tires. To build my intuition for things, I like thinking about them before seeing the explanation, so these are presented with the benchmark information first and the explanation later in case you want to think about your answer before seeing my thoughts. 29. A friend of mine is reviewing performance orders of magnitude to prep for computer performance interviews and found that https://github.com/sirupsen/napkin-math https://github.com/sirupsen/napkin-math 5.4k stars was the top hit. The README's tables include: | Napkin Math performance estimates | | | | | |---|---|---|---|---| | Operation | Latency | Throughput | 1 MiB | 1 GiB | |---|---|---|---|---| | Sequential Memory R/W 64 bytes | 0.5 ns | | | | | ├ Single Thread | | 20 GiB/s | 50 μs | 50 ms | | ├ Threaded | | 200 GiB/s | 5 μs | 5 ms | | Network Same-Zone | | 10 GiB/s | 100 μs | 100 ms | | ├ Inside VPC | | 10 GiB/s | 100 μs | 100 ms | | ├ Outside VPC | | 3 GiB/s | 300 μs | 300 ms | | Hashing, not crypto-safe 64 bytes | 10 ns | 5 GiB/s | 200 μs | 200 ms | | Random Memory R/W 64 bytes | 20 ns | 3 GiB/s | 300 μs | 300 ms | | Fast Serialization 8 9 † | N/A | 1 GiB/s | 1 ms | 1s | | Fast Deserialization 8 9 † | N/A | 1 GiB/s | 1 ms | 1s | | System Call | 300 ns | N/A | N/A | N/A | | Hashing, crypto-safe 64 bytes | 100 ns | 1 GiB/s | 1 ms | 1s | | Sequential SSD read 8 KiB | 1 μs | 8 GiB/s | 100 μs | 100 ms | | Context Switch 1 2 | 10 μs | N/A | N/A | N/A | | Sequential SSD write, -fsync 8KiB | 2 μs | 3 GiB/s | 300 μs | 300 ms | | TCP Echo Server 32 KiB | 50 μs | 500 MiB/s | 2 ms | 2s | | Random SSD Read 8 KiB | 100 μs | 70 MiB/s | 15 ms | 15s | | Decompression 11 | N/A | 1 GiB/s | 1 ms | 1s | | Compression 11 | N/A | 500 MiB/s | 2 ms | 2s | | Sorting 64-bit integers | N/A | 500 MiB/s | 2 ms | 2s | | Proxy: Envoy/ProxySQL/Nginx/HAProxy | 50 μs | ? | ? | ? | | Network within same region | 250 μs | 2 GiB/s | 500 μs | 500 ms | | Premium network within zone/VPC | 250 μs | 25 GiB/s | 50 μs | 40 ms | | Sequential SSD write, +fsync 8KiB | 300 μs | 30 MiB/s | 30 ms | 30s | | {MySQL, Memcached, Redis, ..} Query | 500 μs | ? | ? | ? | | Serialization 8 9 † | N/A | 100 MiB/s | 10 ms | 10s | | Deserialization 8 9 † | N/A | 100 MiB/s | 10 ms | 10s | | Sequential HDD Read 8 KiB | 10 ms | 250 MiB/s | 2 ms | 2s | | Random HDD Read 8 KiB | 10 ms | 0.7 MiB/s | 2 s | 30m | | Blob Storage GET, if-not-match 304 | 30 ms | | | | | Blob Storage GET, 1 conn 128KiB | 80 ms | 100 MiB/s | 10 ms | 10s | | Blob Storage GET, n conn offsets | 80 ms | NW limit | | | | Blob Storage LIST | 100 ms | | | | | Blob Storage PUT, 1 conn 128KiB | 200 ms | 100 MiB/s | 10 ms | 10s | | Blob Storage PUT, n conn multipart | 200 ms | NW limit | 10 ms | 10s | | Network between regions 6 | Varies https://www.cloudping.co/ | 25 MiB/s | 40 ms | 40s | | Network NA Central <- East | 25 ms | 25 MiB/s | 40 ms | 40s | | Network NA Central <- West | 40 ms | 25 MiB/s | 40 ms | 40s | | Network NA East <- West | 60 ms | 25 MiB/s | 40 ms | 40s | | Network EU West <- NA East | 80 ms | 25 MiB/s | 40 ms | 40s | | Network EU West <- NA Central | 100 ms | 25 MiB/s | 40 ms | 40s | | Network NA West <- Singapore | 180 ms | 25 MiB/s | 40 ms | 40s | | Network EU West <- Singapore | 160 ms | 25 MiB/s | 40 ms | 40s | What's wrong with this benchmark? 30. I keep seeing people reference DeepSWE and Senior SWE-Bench to "prove" that their favorite model is better than other people's favorite models or just as generally good benchmarks, such as in What's wrong with these benchmarks? 31. People frequently say that winter tires are superior to all-season tires in cold weather. For example, on googling "all season tires during winter cold" no quotes , the Google AI summary leads with All-season tires lose traction and stiffen in freezing winter temperatures. Their rubber compounds are designed for warmer weather and become hard below 7°C 45°F , leading to significantly longer braking distances and reduced grip ... The rubber in all-season tires cannot maintain pliability in sub-zero temperatures, causing them to perform more like hard plastic on snow and ice. Given that there are a lot of internet comments in the training data, this is a reasonable comment, in that I frequently see variations on this comment on discussions of which tires one should use. What's wrong with this benchmark? 29. Napkin math numbers Random memory access latency The thing that immediately jumped out to my friend Jamie as odd was random memory R/W listed as 20ns, since random memory R/W is implied to be a real DRAM read as opposed to a cache hit , which he felt this should be around 100ns for an order of magnitude estimate. As we were chatting about this, he noted that the README uses the term "latency" for some things that aren't really latencies. Then, when he pulled up the code for random memory read latency, he found the following if you want another exercise, consider what's wrong with the following code before reading the explanation below : while test.i < test.vec.len { let random index = test.order test.i ; black box test.vec random index ; test.i += 1; } Jamie noted that there's no data dependency between the loop iterations, so the memory reads here happen in parallel. Since the alleged latency number is determined by finding the average time for an access, this is incorrect because the CPU can have multiple loads in flight at once. If you wanted to measure latency this way, you'd have to introduce a dependence between loads, to prevent overlapping accesses we discussed a related topic in exercise 19, covered in part 4 of this series https://www.patreon.com/posts/127627543 . Random SSD read I agree with all of Jamie's comments, although I didn't really flag the use of the term latency myself because maybe it's shorthand for latency in some cases and something a bit latency-like in other cases such as reciprocal throughput , which makes the table simpler. What first jumped out to me, besides the memory latency number, was some of the other numbers. For example, random SSD read is listed as 100 us / 70 MB/s. You can get much faster as well as much slower SSDs. For example, if you have a fast but non-exotic, e.g., non Optane device, you might see latencies below 40us, e.g., the Kioxia CD9P-R was measured at ~30 us here https://www.storagereview.com/review/kioxia-cd9p-r-review-read-intensive-gen5-up-to-61-44tb . Other than for some trivial scripts, I haven't worked on anything where I care about disk performance, so I don't have an intuition for what numbers someone would want to have in mind