Conformal Prediction Sets Need Exchangeable Calibration Data
A classifier that emits 0.93 for one class does not automatically provide a statistical statement that the class is correct with probability 0.93. Conformal prediction takes a different route: it uses held-out labeled examples to construct a set of candidate labels with a target marginal coverage level. For split conformal classification, the base model can remain fixed. The guarantee comes from ranking a test example’s conformity or nonconformity score against scores computed on an exchangeable calibration sample. That assumption is the part that gives the coverage statement its scope.