Measure Classifier Calibration Beyond Accuracy
A classifier can keep the same predicted labels while its probability estimates become badly distorted. Accuracy does not expose that change. If a service uses a score of 0.9 to trigger an automated action, the numeric meaning of that score matters independently of whether the top-ranked class is correct. Classifier calibration examines that numeric meaning. For predictions assigned similar confidence, the observed outcome frequency should be close to the stated confidence when the probabilities are well calibrated for the evaluated population.