Skip to content

Archive

Evaluation

7 articles
Artificial Intelligence 16 Sep 2026 5 min read

Measure Classifier Calibration Beyond Accuracy

A classifier can keep the same predicted labels while its probability estimates become badly distorted. Accuracy does not expose that change. If a service uses a score of 0.9 to trigger an automated action, the numeric meaning of that score matters independently of whether the top-ranked class is correct. Classifier calibration examines that numeric meaning. For predictions assigned similar confidence, the observed outcome frequency should be close to the stated confidence when the probabilities are well calibrated for the evaluated population.

Artificial Intelligence 16 Sep 2026 6 min read

Detect Distribution Shift with Energy Scores

A classifier can assign a high softmax probability to an input that does not resemble the data used to fit it. The probability vector still has to sum to one, so normalization can produce a confident-looking prediction even when every class is a poor match. An energy score provides a scalar derived from the logits before that normalization and can serve as a signal for out-of-distribution detection. The score does not make a classifier aware of every possible unfamiliar input. Its value depends on the model, logit scale, training procedure, and data used to set a decision threshold. That makes energy-based detection an evaluation problem as much as a scoring mechanism.

Artificial Intelligence 14 Sep 2026 6 min read

Detect Feature Outliers with Mahalanobis Distance

A feature vector can sit close to a reference mean in Euclidean distance and still be unusual for the distribution that produced the reference data. Mahalanobis distance accounts for this by scaling displacement according to covariance. Directions with little observed variation contribute more to the score than directions in which the reference data naturally spreads out. That behavior makes the distance useful as a compact outlier score for model features or embeddings, provided the reference statistics are meaningful and the covariance estimate is numerically usable.

Artificial Intelligence 13 Sep 2026 5 min read

Calibrate Classifier Confidence with Temperature Scaling

A classifier can choose the correct class yet attach a probability that is too concentrated or too diffuse for the application using that score. Temperature scaling addresses this mismatch after model training by applying one positive scalar to the logits before softmax. The mechanism is deliberately narrow. It changes probability sharpness, not the information represented by the classifier. That boundary makes temperature scaling useful when class ranking is acceptable but confidence values need separate calibration.

Artificial Intelligence 13 Sep 2026 7 min read

Build Classification Sets with Split Conformal Prediction

A classifier normally returns one label or a vector of class scores. Neither output directly states how many labels should remain plausible when the system needs a controlled error rate. Split conformal prediction adds a calibration layer that turns those scores into prediction sets. The useful property is not that every individual set has a fixed probability of containing the correct label. Under the standard exchangeability assumption, split conformal methods target marginal coverage across new examples. That distinction shapes both implementation and interpretation.

Artificial Intelligence 12 Sep 2026 7 min read

Calibrate Neural Classifier Confidence with Temperature Scaling

A neural classifier can choose the correct class often enough for an application while assigning probabilities that are too concentrated or too diffuse. Accuracy alone does not expose this mismatch. A system that acts differently at confidence thresholds also depends on the numerical probabilities attached to its predictions. Temperature scaling is a post-training calibration method that adjusts the sharpness of classifier logits with one positive scalar. For a fixed input, it preserves the ordering of logits, so the predicted class remains unchanged when ordinary argmax decoding is used. What changes is the probability distribution produced after softmax.

Artificial Intelligence 01 Sep 2026 5 min read

Evaluating RAG Systems with a Small Golden Dataset

Retrieval-augmented generation (RAG) is easy to demo and surprisingly hard to evaluate. A fluent answer can hide weak retrieval, while a good retriever can be blamed for an answer model that ignores its evidence. A useful evaluation process separates those failure modes. You do not need thousands of examples to begin. A carefully maintained golden dataset of 30 to 100 representative questions can catch many regressions before users do. Define what the system is supposed to do Start with the product contract rather than a model metric. For a documentation assistant, useful requirements might be: