Evals Are Tests Wearing a Lab Coat
A developer argues that AI evaluations are structurally identical to traditional software tests, just rebranded with new terminology. The post claims that the eval-tooling stack—golden datasets, judges, scorers, rubrics—…