Skip to content

Archive

Classification

28 articles
Artificial Intelligence 24 Sep 2026 5 min read

Temperature Scaling Recalibrates Classifier Confidence Without Changing Class Order

A classifier can rank the correct class above every alternative yet attach probabilities that are systematically too concentrated or too diffuse. Temperature scaling addresses that mismatch after training by applying one scalar to the logits before softmax. It changes reported confidence without changing the underlying classifier parameters. The mechanism is narrow. It does not repair incorrect class rankings, add information to the representation, or make every individual probability accurate. Its target is the relationship between confidence and observed outcomes on data representative of deployment.

Artificial Intelligence 24 Sep 2026 4 min read

Label Smoothing Redistributes Target Probability Across Classes

A classifier trained with one-hot targets assigns all target probability mass to one class. Label smoothing changes that target before cross-entropy is evaluated: some mass is moved away from the designated class and assigned to other classes. The network architecture can remain identical, yet the optimization objective is no longer the same. That distinction matters when interpreting confidence, loss values, and implementation settings. Label smoothing is not a post-processing operation on predicted probabilities. It changes the target distribution used to produce the training signal.

Artificial Intelligence 23 Sep 2026 5 min read

Label Smoothing Redistributes Target Probability Before Cross-Entropy

A classifier trained with one-hot targets is asked to place all target probability on a single class. Cross-entropy does not require that target representation. Label smoothing changes the target distribution before the loss is evaluated, so the model receives a different gradient even when its logits and predicted probabilities are unchanged. This distinction matters because label smoothing is not a decoding rule and does not alter inference by itself. It changes the training objective. The resulting model parameters can differ because the optimizer follows gradients computed against softened targets.

Artificial Intelligence 17 Sep 2026 6 min read

Use Selective Classification to Trade Coverage for Error Rate

A classifier usually returns a label for every input, even when its score distribution is nearly tied or the input sits far from familiar data. Selective classification changes that interface: the system may return a prediction or abstain. The acceptance rule then determines both how many inputs receive predictions and how often those accepted predictions are wrong. This is not the same as making the classifier intrinsically more accurate. Abstention moves some cases out of the automatic-decision set. Its value depends on whether the selection score ranks difficult cases well enough for rejected inputs to contain a disproportionate share of errors.

Artificial Intelligence 17 Sep 2026 5 min read

Evaluate Probabilistic Classifiers with the Brier Score

Two classifiers can produce the same predicted labels and the same accuracy while assigning very different probabilities to those labels. A system that emits 0.51 for every correct binary decision is not making the same probabilistic claim as one that emits 0.99, even though thresholded accuracy may treat them identically. The Brier score keeps that distinction visible. It measures squared error between predicted probabilities and observed outcomes, so both the selected class and the probability assigned to each outcome affect the result. This makes it useful when downstream code consumes probabilities for ranking, thresholds, abstention, or expected-cost decisions.

Artificial Intelligence 17 Sep 2026 6 min read

Conformal Prediction Sets Need Exchangeable Calibration Data

A classifier that emits 0.93 for one class does not automatically provide a statistical statement that the class is correct with probability 0.93. Conformal prediction takes a different route: it uses held-out labeled examples to construct a set of candidate labels with a target marginal coverage level. For split conformal classification, the base model can remain fixed. The guarantee comes from ranking a test example’s conformity or nonconformity score against scores computed on an exchangeable calibration sample. That assumption is the part that gives the coverage statement its scope.

Artificial Intelligence 16 Sep 2026 5 min read

Measure Classifier Calibration Beyond Accuracy

A classifier can keep the same predicted labels while its probability estimates become badly distorted. Accuracy does not expose that change. If a service uses a score of 0.9 to trigger an automated action, the numeric meaning of that score matters independently of whether the top-ranked class is correct. Classifier calibration examines that numeric meaning. For predictions assigned similar confidence, the observed outcome frequency should be close to the stated confidence when the probabilities are well calibrated for the evaluated population.

Artificial Intelligence 16 Sep 2026 6 min read

Detect Distribution Shift with Energy Scores

A classifier can assign a high softmax probability to an input that does not resemble the data used to fit it. The probability vector still has to sum to one, so normalization can produce a confident-looking prediction even when every class is a poor match. An energy score provides a scalar derived from the logits before that normalization and can serve as a signal for out-of-distribution detection. The score does not make a classifier aware of every possible unfamiliar input. Its value depends on the model, logit scale, training procedure, and data used to set a decision threshold. That makes energy-based detection an evaluation problem as much as a scoring mechanism.

Artificial Intelligence 15 Sep 2026 6 min read

Account for Label Smoothing in Classifier Confidence

A classifier trained with one-hot targets is rewarded for moving probability mass toward the labeled class. Cross-entropy keeps decreasing as the model assigns that class a probability closer to one, even after the predicted class is already correct. Label smoothing changes this pressure by assigning a small amount of target mass to the other classes. That change is easy to treat as a minor detail in the loss function. It is not minor when an application consumes the model’s probability values. The smoothed target changes the optimum encouraged by the training objective, so confidence scores from a smoothed model should not be interpreted as if they came from the same objective as ordinary one-hot training.

Artificial Intelligence 14 Sep 2026 7 min read

Control Target Certainty with Label Smoothing

A classifier trained with ordinary cross-entropy often receives a one-hot target: probability mass 1 on the labeled class and 0 on every other class. That target keeps rewarding movement toward a more extreme prediction even after the correct class already has the highest score. Label smoothing changes the target distribution rather than the model architecture. A small amount of target mass is moved away from the labeled class and assigned to other classes. Cross-entropy then optimizes against this softened distribution, so the gradient no longer treats absolute certainty on the labeled class as the target state.

Artificial Intelligence 14 Sep 2026 5 min read

Control Classifier Logits with Cosine Normalization

A linear classification head mixes two signals in each logit: the angle between a feature vector and a class weight vector, and the magnitudes of both vectors. Cosine normalization removes the magnitude terms, so class scores depend on directional alignment instead. That change is small in code but substantial in interpretation. Feature norm no longer increases every class comparison merely by growing, class-weight norm no longer acts as an implicit class-specific scale, and the overall sharpness of the softmax must be supplied separately.

Artificial Intelligence 13 Sep 2026 6 min read

Regularize Classifier Targets with Label Smoothing

A classifier trained with one-hot targets is rewarded for pushing the target class probability toward one and every other class probability toward zero. Cross-entropy keeps applying pressure in that direction even after the predicted class is already correct. Label smoothing changes that pressure by replacing the exact one-hot target with a distribution that reserves some mass for other classes. That small change affects more than the target tensor. It changes the gradient on every output logit, limits the incentive for extreme class separation, and alters how predicted probabilities should be interpreted.

Artificial Intelligence 13 Sep 2026 5 min read

Focus Classification Loss with Focal Modulation

Cross-entropy gives every classified example a loss determined by the probability assigned to its target class. When a training batch contains many examples the model already classifies with high confidence, their individual losses may be small yet their aggregate contribution can still occupy a substantial part of the objective. Focal loss changes that balance with a confidence-dependent multiplier. The mechanism is not a new classifier head or sampling strategy. It modifies the loss so that examples with high target-class probability are attenuated more strongly than examples with low target-class probability.

Artificial Intelligence 13 Sep 2026 5 min read

Calibrate Classifier Confidence with Temperature Scaling

A classifier can choose the correct class yet attach a probability that is too concentrated or too diffuse for the application using that score. Temperature scaling addresses this mismatch after model training by applying one positive scalar to the logits before softmax. The mechanism is deliberately narrow. It changes probability sharpness, not the information represented by the classifier. That boundary makes temperature scaling useful when class ranking is acceptable but confidence values need separate calibration.

Artificial Intelligence 13 Sep 2026 7 min read

Build Classification Sets with Split Conformal Prediction

A classifier normally returns one label or a vector of class scores. Neither output directly states how many labels should remain plausible when the system needs a controlled error rate. Split conformal prediction adds a calibration layer that turns those scores into prediction sets. The useful property is not that every individual set has a fixed probability of containing the correct label. Under the standard exchangeability assumption, split conformal methods target marginal coverage across new examples. That distinction shapes both implementation and interpretation.

Artificial Intelligence 12 Sep 2026 7 min read

Soften Classification Targets with Label Smoothing

A classifier trained with one-hot targets receives a strong signal to push the target class probability toward one and every other class probability toward zero. Cross-entropy supports that behavior even after the predicted class is already correct: making the target probability more extreme can still reduce the loss. Label smoothing changes the target distribution before cross-entropy is computed. Instead of assigning all target mass to one class, it reserves a small amount for the remaining classes. This alters the gradient applied to the logits and reduces pressure toward extreme output distributions.

Artificial Intelligence 11 Sep 2026 11 min read

Reweight Long-Tailed Classification with Effective Sample Counts

A classifier trained on a long-tailed dataset can see thousands of examples from common classes and only a handful from rare ones. Ordinary empirical risk minimization gives the common classes more influence simply because they appear more often. A tempting fix is to weight each class by the inverse of its example count, but that can make a tiny class disproportionately influential, including any mislabeled examples it contains. Class-balanced loss based on the effective number of samples provides a smoother way to derive class weights. Instead of treating every additional example as equally informative, it models diminishing returns within a class and weights classes according to an adjusted, or effective, sample count.

Artificial Intelligence 11 Sep 2026 9 min read

Focus Classifier Training with Focal Loss

Focus Classifier Training with Focal Loss A classifier can spend much of its training signal on examples it already handles correctly. This becomes especially troublesome when easy examples vastly outnumber difficult ones. A detector, for instance, may encounter many obvious background locations for every location containing an object. Focal loss changes the contribution of each example according to the model’s confidence in the correct class. Easy, high-confidence examples receive less weight. Harder examples retain more of their cross-entropy loss. The mechanism is small, but using it well requires understanding what it changes and what it doesn’t.

Artificial Intelligence 10 Sep 2026 10 min read

Neural Collapse in Deep Classifiers

Neural Collapse in Deep Classifiers A classifier can keep changing after it already predicts every training example correctly. Cross-entropy loss can continue to fall, feature vectors can reorganize, and the final classification layer can become increasingly regular. Looking only at training accuracy hides all of that movement. Neural collapse is a name for a collection of geometric patterns that can emerge late in the training of deep classifiers. The striking part isn’t simply that examples from the same class become similar. Under the conditions where neural collapse appears, within-class variation can shrink while class centers and classifier weights approach a highly symmetric arrangement.

Artificial Intelligence 07 Sep 2026 9 min read

Defer Uncertain Classifier Predictions with Selective Classification

A classifier does not have to make an automated decision for every input. In many systems, forcing a prediction on the hardest cases is exactly what creates expensive mistakes. Imagine a model that routes customer support messages. Clear password-reset requests can be handled automatically, while ambiguous messages could be sent to a human queue. The important design question is no longer only “How accurate is the classifier?” It is also “How accurate is it on the cases we allow it to handle?”

Artificial Intelligence 06 Sep 2026 12 min read

Cross-Entropy Loss for Classification

A classifier needs more than a way to count correct answers. During training, it needs a signal that says not only whether a prediction was wrong, but also how the model’s scores should change. Suppose the correct class is cat. A model that assigns cat probability 0.49 and another class 0.51 is wrong, but it is close to the decision boundary. A model that assigns cat probability 0.001 is also wrong, and much more confident in that mistake. Treating those predictions as equally bad throws away useful information.

Artificial Intelligence 05 Sep 2026 9 min read

Use Class-Weighted Loss for Imbalanced Classification

A classifier trained on imbalanced data can achieve a low average loss while learning the minority class poorly. If 99% of training examples belong to one class, errors on the remaining 1% contribute relatively little to an unweighted objective simply because they occur less often. Class-weighted loss changes that training signal. Instead of treating every example’s loss equally, it gives examples from selected classes more influence on parameter updates. This is useful when class frequency and the importance of learning each class are badly misaligned.

Artificial Intelligence 04 Sep 2026 10 min read

Macro F1 and Balanced Accuracy for Imbalanced Classifiers

A classifier can report impressive accuracy while failing on the cases you care about most. This happens easily when one class is much more common than another. Imagine a model that detects defective components on a production line. In a test set of 1,000 components, 950 are normal and 50 are defective. A model that predicts normal for every component is correct 95% of the time, yet it detects none of the defects.

Artificial Intelligence 04 Sep 2026 9 min read

Handle Class Imbalance in Machine Learning

A classifier can achieve impressive accuracy while failing on the cases you care about most. If only 1% of transactions are fraudulent, a model that predicts “not fraud” for every transaction is 99% accurate and still useless for detecting fraud. This is the practical problem of class imbalance: some target classes appear much less often than others. Imbalance does not automatically make a dataset bad, and it does not imply that every model needs special treatment. It does mean that accuracy can hide important errors and that the training objective may give rare examples too little influence.