Skip to content

Archive

Machine Learning

98 articles
Artificial Intelligence 08 Sep 2026 9 min read

Regularize Residual Networks with Stochastic Depth

Deep residual networks can overfit even when their skip connections make optimization manageable. Standard dropout can regularize individual activations, but residual architectures offer another useful unit to randomize: the entire residual branch. Stochastic depth randomly removes selected residual branches during training while keeping the skip path intact. A training example may therefore pass through a slightly shallower effective network on one step and the full set of blocks on another. At inference time, every residual branch is normally active.

Artificial Intelligence 08 Sep 2026 10 min read

Model Soups for Combining Fine-Tuned Models

A hyperparameter sweep often leaves you with several fine-tuned models that are individually useful. The usual workflow keeps the checkpoint with the best validation score and discards the rest. An ensemble can use several checkpoints, but then every request may require multiple model evaluations, increasing inference cost and operational complexity. A model soup offers a third option: average the parameters of compatible fine-tuned models and deploy the resulting parameter set as one model. The technique is simple, but its simplicity can be misleading. Parameter averaging is meaningful only when the checkpoints are sufficiently compatible, and the averaged model still needs independent evaluation.

Artificial Intelligence 08 Sep 2026 9 min read

Load Balancing in Mixture-of-Experts Models

A mixture-of-experts model can contain many expert networks while activating only a small subset for each token. That sparse computation is attractive because the model can have more parameters without evaluating every parameter for every token. But sparsity creates a new problem: the router can send too many tokens to the same experts. If one expert receives most of a batch while others sit nearly idle, the model does not get the practical benefit that its expert count suggests. In systems with fixed expert capacity, overloaded experts can also overflow, so some token-to-expert assignments cannot be processed as intended.

Artificial Intelligence 08 Sep 2026 10 min read

Dynamical Isometry in Neural Networks

A deep neural network can look well scaled one layer at a time and still be difficult to optimize. Signals pass through many transformations, and small expansions or contractions can multiply with depth. By the time a gradient travels through the whole network, some directions may have nearly disappeared while others have been amplified dramatically. Dynamical isometry gives a precise way to reason about this problem. Instead of asking only whether the average gradient magnitude is reasonable, it asks how the network transforms different directions in its input space. The relevant object is the input-output Jacobian, and its singular values provide the main measurements.

Artificial Intelligence 08 Sep 2026 9 min read

Detect Out-of-Distribution Inputs with Energy Scores

A classifier can be highly accurate on its test set and still behave confidently on inputs that are unlike anything it was trained to recognize. A product classifier trained on shoes, bags, and watches may receive a photo of a bicycle and still be forced to choose one of its known classes. That creates a deployment problem: ordinary classification answers which known class looks most likely, but many systems also need to ask whether this input resembles the data on which the classifier was validated.

Artificial Intelligence 08 Sep 2026 8 min read

Deep Ensembles for Model Uncertainty

A neural network can return a confident prediction even when an input is unfamiliar or ambiguous. Looking only at one model’s largest probability can therefore hide an important question: would another plausible model, trained on the same task, make the same decision? A deep ensemble helps answer that question by training several neural networks independently and combining their predictions. The combined prediction can improve robustness in some settings, while disagreement among members provides a practical uncertainty signal. It is not a guarantee that the prediction is correct, and it does not detect every kind of uncertainty.

Artificial Intelligence 07 Sep 2026 12 min read

Xavier and He Initialization for Neural Networks

A deep neural network can fail before learning has had a fair chance. If its initial weights make activations or gradients shrink layer after layer, useful signals can become tiny. If those quantities grow too much, training can become unstable. The optimizer may receive a problem that is unnecessarily difficult even though the architecture and data are otherwise reasonable. Weight initialization tries to start the network in a numerically useful regime. Two common schemes are Xavier initialization, also called Glorot initialization, and He initialization, also called Kaiming initialization. Both choose the scale of random weights from the size of a layer, but they make different assumptions about how signals pass through the activation function.

Artificial Intelligence 07 Sep 2026 9 min read

Weight Decay in Adam and AdamW

A training configuration can contain a parameter named weight_decay without making it obvious what operation the optimizer actually performs. That ambiguity matters most with adaptive optimizers such as Adam: adding an L2 penalty to the loss and directly decaying parameters are not generally the same update. The distinction is easy to miss because the two procedures are closely related under ordinary stochastic gradient descent (SGD). Once an optimizer rescales different coordinates using gradient history, however, the equivalence breaks.

Artificial Intelligence 07 Sep 2026 9 min read

Transfer Hyperparameters Across Model Width with MuP

Scaling a neural network creates an expensive tuning problem. A learning rate that works for a small prototype may behave differently after hidden dimensions become much wider. If every model size needs a fresh hyperparameter sweep, experimenting on small models saves less compute than it first appears. Maximal Update Parametrization, usually written MuP or μP, addresses this problem by changing how parameter initialization and learning rates scale with model width. The goal is not to make a wider model identical to a narrow one. It is to make important training dynamics behave consistently enough that hyperparameters tuned on a smaller proxy can often transfer to a wider target.

Artificial Intelligence 07 Sep 2026 10 min read

Train Neural Networks with Sharpness-Aware Minimization

A neural network can reach low training loss at parameter values where a small change to the weights makes the loss rise sharply. Standard optimization does not directly discourage this behavior: it mainly asks whether the loss is low at the current parameters. Sharpness-Aware Minimization (SAM) changes the training objective. Instead of optimizing only the loss at the current weights, it approximately optimizes the worst loss in a small neighborhood around them. The practical idea is simple: find a nearby parameter perturbation that makes the current mini-batch harder, then update the original model using the gradient measured at those perturbed parameters.

Artificial Intelligence 07 Sep 2026 10 min read

Resolve Conflicting Gradients in Multi-Task Learning

Training one neural network to solve several tasks can reduce duplicated computation and let related tasks share useful representations. It also creates a problem that single-task training does not have: two losses can ask the same shared parameter to move in opposing directions during the same update. Simply adding the losses does not make that disagreement disappear. Their gradients are added too, so one task can partially cancel another or dominate the shared update. Gradient surgery is a family of techniques that changes task gradients before combining them. A well-known example is projected conflicting gradients, commonly called PCGrad, which removes a conflicting component of one task’s gradient relative to another.

Artificial Intelligence 07 Sep 2026 10 min read

Probe Neural Network Representations with Linear Classifiers

A neural network can produce the right output while leaving an important engineering question unanswered: what information exists inside its intermediate representations? Suppose an image classifier predicts product categories. You may want to know whether an early layer already separates shapes, whether a later layer distinguishes categories, or whether a supposedly irrelevant attribute such as camera source remains easy to recover. Looking only at the final prediction does not answer those questions.

Artificial Intelligence 07 Sep 2026 9 min read

Prevent Catastrophic Forgetting in Continual Learning

Updating a neural network with new data sounds straightforward: continue training on the new examples and deploy the improved model. The difficulty is that an update which helps the new data can damage behavior the model learned earlier. A classifier that learns a new group of products, for example, may become worse at recognizing older groups even though those old classes never changed. This failure is called catastrophic forgetting. It matters most in continual learning, where a model learns from a sequence of tasks or data distributions instead of training once on a fixed mixed dataset.

Artificial Intelligence 07 Sep 2026 9 min read

Knowledge Distillation for Smaller Classifiers

A model can be accurate enough for a product and still be too expensive to deploy. A large classifier may exceed a mobile memory budget, miss a latency target, or cost too much when every request requires substantial compute. Replacing it with a smaller model reduces those costs, but training the smaller model only from ground-truth labels can leave useful information behind. Knowledge distillation addresses this problem by training a smaller student model to learn from a stronger teacher model. Instead of seeing only the correct class, the student can also learn how the teacher distributes its confidence across the alternatives.

Artificial Intelligence 07 Sep 2026 13 min read

Handle Label Shift in Deployed Classifiers

A classifier can keep seeing familiar inputs and still make worse decisions after deployment. One reason is that the frequency of the classes has changed. Imagine a model trained to classify support tickets as billing, account, or technical. During training, billing tickets made up 20% of examples. After a pricing migration, billing issues temporarily rise to 50%. The model has not changed, but one part of the environment has: the prior probability of each class.

Artificial Intelligence 07 Sep 2026 12 min read

Gradient Noise Scale and Batch Size

Increasing a neural network’s batch size can make more accelerators useful, but the benefit does not grow indefinitely. At some point, processing more examples before each update gives a cleaner estimate of nearly the same gradient direction while consuming additional examples and compute. Gradient noise scale provides a useful mental model for this transition. It compares the variation in per-example gradients with the strength of their average. When gradient estimates are noisy relative to their mean, averaging more examples can remove meaningful noise. When they are already stable, a larger batch has less statistical work left to do.

Artificial Intelligence 07 Sep 2026 9 min read

Gradient Centralization in Neural Network Training

Neural network optimizers normally consume the gradients produced by backpropagation directly. Gradient centralization inserts one small transformation between those two steps: for selected weight tensors, it subtracts the mean of each gradient vector before the optimizer uses it. That operation is easy to implement, but its effect is easy to misunderstand. It is not gradient clipping, because it does not cap large values. It is not normalization, because it does not divide by a norm or standard deviation. It changes the direction of the update by removing one particular component.

Artificial Intelligence 07 Sep 2026 10 min read

Diagnose and Prevent Dead ReLU Neurons

ReLU is a simple and effective activation function, but its simplicity creates a failure mode that is easy to miss. A neuron can reach a state where its pre-activation is negative for every relevant input. Its ReLU output is then always zero, and the gradient through that activation is also zero. If this persists, the neuron may stop participating in learning. This is commonly called a dead ReLU or dying ReLU problem. It does not mean that every zero activation is a defect: sparse activations are a normal consequence of ReLU. The useful question is whether a unit is inactive only for some inputs or effectively inactive across the data it needs to model.

Artificial Intelligence 07 Sep 2026 9 min read

Defer Uncertain Classifier Predictions with Selective Classification

A classifier does not have to make an automated decision for every input. In many systems, forcing a prediction on the hardest cases is exactly what creates expensive mistakes. Imagine a model that routes customer support messages. Clear password-reset requests can be handled automatically, while ambiguous messages could be sent to a human queue. The important design question is no longer only “How accurate is the classifier?” It is also “How accurate is it on the cases we allow it to handle?”

Artificial Intelligence 07 Sep 2026 11 min read

Compress Neural Network Layers with Low-Rank Factorization

Large neural networks spend much of their memory and computation multiplying activations by weight matrices. Some of those matrices contain more independent structure than the model actually needs for a particular deployment. If so, we can approximate one large matrix with two smaller matrices and reduce the number of stored parameters and multiply-add operations. This technique is called low-rank factorization. The central idea is simple, but using it well requires more than choosing a smaller number. Compression changes the weights, approximation error can accumulate through a network, and fewer arithmetic operations do not guarantee lower wall-clock latency on every device.

Artificial Intelligence 07 Sep 2026 10 min read

Adapt Neural Networks with Gradient Reversal

A classifier can perform well in evaluation and then weaken after deployment because the inputs changed. Product photos may come from a new camera, support messages may use different vocabulary, or sensor readings may come from different hardware. Collecting labels for the new environment can be expensive even when unlabeled examples are easy to obtain. Domain-adversarial training addresses one version of this problem. It asks a feature extractor to support the prediction task while making the source and target domains difficult to distinguish. A gradient reversal layer makes those two goals trainable with ordinary backpropagation by reversing the domain classifier’s gradient before it reaches the feature extractor.

Artificial Intelligence 06 Sep 2026 10 min read

Stabilize Neural Network Evaluation with Exponential Moving Average Weights

A neural network’s final training step is not necessarily its most useful checkpoint. Stochastic optimization keeps moving the parameters as it follows noisy mini-batch gradients, so two nearby checkpoints can behave slightly differently even when training is otherwise healthy. An exponential moving average (EMA) of model weights gives you a second set of parameters that changes more smoothly. Instead of evaluating only the latest training weights, you maintain a weighted history in which recent weights matter most and older weights gradually fade away.

Artificial Intelligence 06 Sep 2026 11 min read

Reduce Transformer Inference Cost with Early Exiting

A transformer classifier normally spends the same number of layers on every input. A clear support request and an ambiguous one both pass through the entire network, even when an intermediate representation already contains enough information to classify the easy case correctly. Early exiting changes that fixed-compute rule. It attaches prediction heads to intermediate layers and lets an input stop once a chosen exit rule considers the prediction sufficiently reliable. Easy inputs can use less computation, while harder inputs continue through deeper layers.

Artificial Intelligence 06 Sep 2026 12 min read

Cross-Entropy Loss for Classification

A classifier needs more than a way to count correct answers. During training, it needs a signal that says not only whether a prediction was wrong, but also how the model’s scores should change. Suppose the correct class is cat. A model that assigns cat probability 0.49 and another class 0.51 is wrong, but it is close to the decision boundary. A model that assigns cat probability 0.001 is also wrong, and much more confident in that mistake. Treating those predictions as equally bad throws away useful information.