Holdout Ledgers Keep Agent Scores Honest
A developer has proposed a methodology for keeping coding-agent benchmark scores reproducible by treating evaluation datasets as versioned, hash-pinned artifacts rather than mutable files. The approach requires task reco…