What Happens When AI Outgrows the Tests We Use to Measure It?
A developer argues that benchmark saturation and shifting evaluation methods make it increasingly difficult to judge how much better new AI models like GPT-6 Astra actually are, citing Astra's system card examples of ret…