Skip to content

Archive

Data Science

6 articles
Data Science 02 Sep 2026 5 min read

Robust Outlier Detection with Median Absolute Deviation

Outlier detection often begins with a rule such as “flag values more than three standard deviations from the mean.” That works reasonably well for approximately normal data without severe contamination, but the same extreme observations you want to detect can move both the mean and the standard deviation. Median absolute deviation, usually abbreviated MAD, provides a more robust alternative. Why mean and standard deviation can be fragile Consider response times in milliseconds:

Data Science 02 Sep 2026 5 min read

Bootstrap Confidence Intervals with Resampling

A point estimate hides uncertainty. Reporting that median latency is 180 ms or a conversion-rate difference is 1.4 percentage points does not show how much that estimate might move if another sample were collected. Bootstrap resampling is a practical way to estimate sampling uncertainty when deriving an analytic formula is difficult or when the statistic is not a simple mean. The bootstrap idea Given an observed sample of size n:

Data Science 01 Sep 2026 4 min read

Time Series Cross-Validation with Walk-Forward Splits

Random train/test splits assume examples are exchangeable. Time-series data violates that assumption because the future occurs after the past, and production models normally predict observations that were not available during training. Walk-forward validation preserves that chronology. Why random splitting is misleading Suppose you want to predict next week’s demand from historical sales. A random split can place March observations in the test set while April observations appear in training. Even if features do not explicitly contain future values, the evaluation now uses a model fitted on a future regime. Seasonality, pricing, inventory, customer behavior, and economic conditions can all make the score more optimistic than deployment reality.

Data Science 01 Sep 2026 5 min read

Probability Calibration for Classification Models

A classifier can rank examples correctly while producing probabilities that are poor estimates of real-world likelihood. If a model assigns 0.8 probability to many comparable cases, calibration asks whether roughly 80% of those cases are actually positive. This matters whenever probabilities drive decisions such as pricing, triage, alert thresholds, expected value, or human review. Discrimination and calibration are different Metrics such as ROC AUC evaluate how well a model ranks positive examples above negative ones. They do not require predicted probabilities to match observed frequencies.

Data Science 01 Sep 2026 3 min read

Calibrate Classification Probabilities Before Using Decision Thresholds

A classifier can rank examples well while producing poor probability estimates. If predictions drive cost-sensitive decisions, triage, or risk thresholds, the difference matters. A score of 0.8 is useful as a probability only when similarly scored examples are positive about 80% of the time under the deployment distribution. Separate discrimination from calibration Metrics such as ROC AUC primarily measure ranking. Calibration asks whether predicted probabilities agree with observed frequencies. A model can have strong AUC and still be overconfident or underconfident.

Data Science 01 Sep 2026 5 min read

Avoiding Data Leakage in Machine Learning Pipelines

Data leakage happens when information that would not be available at prediction time influences model training. The result is an evaluation score that looks excellent in development and collapses after deployment. Leakage is often subtle because the model code itself can be correct. The mistake lives in how datasets, features, preprocessing, and time boundaries are constructed. Split before learning from the data A classic mistake is standardizing the full dataset and splitting afterward.