AI Future Leakage: The Silent Flaw Breaking How We Test Whether Machines Can Predict the Future Northwestern University researchers have demonstrated that the standard method for detecting training-data contamination in AI forecasting benchmarks is mathematically flawed, after a test they built falsely accused four of five flagship AI models of cheating. The test used questions resolved only after each model's training cutoff, meaning the answers could not have been memorized, yet the conventional contamination check still flagged the models as contaminated. Four of five flagship AI models just failed a contamination test they could not possibly have failed. Every question they were scored on was resolved after their training data ended, meaning none of the answers could have been memorized. Northwestern researchers built the test anyway, watched it accuse innocent models, and proved mathematically why the entire field’s standard method for catching this problem has been broken all along