Does Your AI Know When It Might Be Wrong?
A model evaluated by an unnamed developer scored 88% accuracy on a four-stage clinical staging task, but 44% of its mistakes were made while it was more than 80% confident, highlighting a calibration gap that accuracy ca…