Four of five flagship AI models just failed a contamination test they could not possibly have failed. Every question they were scored on was resolved after their training data ended, meaning none of the answers could have been memorized. Northwestern researchers built the test anyway, watched it accuse innocent models, and proved mathematically why the entire field’s standard method for catching this problem has been broken all along
Microsoft AI Chief Says China Isn’t Excuse to Forego Regulation