Fast Counterfactuals for Go Cache Hit Rates CloudX reported that 86% of actions/setup-go test package runs are pre-empted by its improved GitHub Actions caching strategy, according to the company's post on scaling Go CI. Because the company never benchmarked cloudx-io/setup-go against actions/setup-go simultaneously, it backtested the claim over a 4,000-commit range using go list build IDs to detect test package changes without running tests. CloudX said replaying or simulating 4,000 commits was infeasible because running the suite at an optimistic two minutes per commit would take more than five days per tested treatment. The pièce de résistance in ‘Scaling Golang CI by Replacing actions/setup-go ’ https://www.cloudx.ai/posts/setup-go is the conclusion that 86% of actions/setup-go test package runs are pre-empted by our improved caching strategy. To summarize the CloudX article, the puzzle is to optimize use of a fine-grain cache we don’t control the GOCACHE which is saved and loaded through a coarse-grain cache where we can control cache keys that determine which instances of the GOCACHE are loaded and saved. CloudX makes that outer, coarse-grained GitHub Actions cache more granular by writing to it often and restoring fresher entries. This comparison between our key strategy and GitHub’s is a counterfactual comparison: the course of action we took against a hypothetical alternative. These are tricky We didn’t foresee open-sourcing cloudx-io/setup-go , so we never ran it alongside actions/setup-go simultaneously to benchmark their relative performance; we just switched, realized savings, and never looked back. To claim a general improvement — not just an improvement for a short, potentially unrepresentative period — I needed to show an advantage over an extended interval, including different paces of development on different kinds of features. I picked a 4,000-commit range, our latest few months of development. How would you backtest GitHub Actions performance for 4,000 commits? Maybe replay comes to mind. We have the full commit history; we could create a repository using each Actions strategy and apply each monorepo commit one at a time, recording cache-hit rates at each commit. Alternatively, we could simulate the Actions behavior locally and using the GOCACHE environment variable to point go test runs at the faux-Actions cache to “restore.” Neither of these approaches works because running tests is slow. Even if the suite takes two minutes to run per commit on average optimistic , testing each of 4,000 commits in series would take more than five days per tested treatment.