Skip to content

Archive

Machine Learning

98 articles
Artificial Intelligence 11 Sep 2026 9 min read

Use Token Dropout to Train Robust Sequence Models

Use Token Dropout to Train Robust Sequence Models A sequence model can become too dependent on a few easy input clues. Remove one field, truncate a message, or corrupt a token at inference time, and a prediction that looked reliable on clean validation data may change sharply. Token dropout is a simple training-time corruption technique: randomly hide some input tokens while keeping the learning target unchanged. The model is forced to solve some training examples without every usual clue. Used carefully, this can reduce brittle dependence on individual tokens. Used carelessly, it can destroy information the task genuinely requires.

Artificial Intelligence 11 Sep 2026 11 min read

Reweight Long-Tailed Classification with Effective Sample Counts

A classifier trained on a long-tailed dataset can see thousands of examples from common classes and only a handful from rare ones. Ordinary empirical risk minimization gives the common classes more influence simply because they appear more often. A tempting fix is to weight each class by the inverse of its example count, but that can make a tiny class disproportionately influential, including any mislabeled examples it contains. Class-balanced loss based on the effective number of samples provides a smoother way to derive class weights. Instead of treating every additional example as equally informative, it models diminishing returns within a class and weights classes according to an adjusted, or effective, sample count.

Artificial Intelligence 11 Sep 2026 11 min read

Prevent VAE Posterior Collapse with Free Bits

Prevent VAE Posterior Collapse with Free Bits A variational autoencoder can appear to train normally while its latent representation becomes nearly useless. The decoder learns to explain the data without depending on the latent variable, the encoder moves toward the prior, and the KL divergence shrinks toward zero. This failure mode is called posterior collapse. Free bits is a small change to the VAE objective that can reduce one source of that collapse. It stops the KL term from rewarding the optimizer for squeezing an already-small amount of latent information even closer to zero. The technique is simple, but its name and common shorthand can lead to a misleading mental model. Free bits does not force a latent variable to contain a chosen amount of information. It changes the optimization pressure below a threshold.

Artificial Intelligence 11 Sep 2026 11 min read

Let Classifiers Abstain with Selective Prediction

Let Classifiers Abstain with Selective Prediction A classifier does not have to answer every case. In many systems, forcing a prediction on an ambiguous input is worse than sending that input to a human, requesting more information, or falling back to a safer workflow. Selective prediction gives a model that option. The classifier produces its usual prediction, but the system accepts it only when a selection rule considers the case reliable enough. Otherwise, the system abstains.

Artificial Intelligence 10 Sep 2026 9 min read

Regularize Neural Networks with Mixup

Regularize Neural Networks with Mixup A neural network can fit its training examples while behaving unpredictably in the space between them. If two nearby inputs belong to different classes, standard training tells the model what to do at the endpoints but often says little about intermediate points. Mixup changes that training signal. Instead of training only on individual examples, it creates synthetic examples by interpolating pairs of inputs and their labels. The model is then asked to make a correspondingly mixed prediction. This acts as a regularizer because it constrains how predictions may change between training examples.

Artificial Intelligence 10 Sep 2026 9 min read

Reduce Repetitive Generation with Unlikelihood Training

Reduce Repetitive Generation with Unlikelihood Training A language model can learn to predict ordinary text well and still assign too much probability to behavior you don’t want at generation time. Repetition is a common example: once a phrase appears, the model may keep making recently used tokens plausible enough that a decoding loop becomes hard to escape. Changing the decoder can hide some of this behavior, but it doesn’t change the probabilities learned by the model. Unlikelihood training takes a different approach. During training, it identifies undesirable candidates and explicitly pushes their probabilities down while the usual likelihood objective pushes desired tokens up.

Artificial Intelligence 10 Sep 2026 10 min read

Prevent FP16 Gradient Underflow with Dynamic Loss Scaling

Prevent FP16 Gradient Underflow with Dynamic Loss Scaling Mixed-precision training can reduce memory use and accelerate supported operations, but float16 introduces a numerical problem that is easy to miss: some gradients are too small to survive in FP16. They can round to zero before the optimizer gets a chance to use them. Loss scaling addresses that problem by multiplying the loss before backpropagation, which multiplies the resulting gradients by the same factor. The gradients are divided by that factor before the optimizer update, so the intended update is unchanged when the arithmetic remains finite. Dynamic loss scaling adjusts the factor during training so you don’t have to guess one fixed value for the whole run.

Artificial Intelligence 10 Sep 2026 9 min read

Prepare Models for Low-Precision Inference with Quantization-Aware Training

A model can work well in floating point and lose useful accuracy after its weights or activations are quantized for deployment. The problem is not mysterious: rounding and clipping change the numbers that flow through the network, while the original model was optimized without those changes in the loop. Quantization-aware training (QAT) exposes the model to an approximation of those low-precision numerics while its parameters can still adapt. Training remains differentiable in floating point, but the forward computation simulates the quantization errors expected after conversion.

Artificial Intelligence 10 Sep 2026 10 min read

Distill Sequence Models with Teacher-Generated Outputs

Distill Sequence Models with Teacher-Generated Outputs A large text generator may produce useful outputs but still be too expensive for the latency, memory, or throughput budget of a deployment. Training a smaller model on the original dataset is the obvious baseline, but it throws away information encoded in the larger model’s behavior. Sequence-level knowledge distillation offers another option: let a capable teacher generate target sequences, then train a smaller student to reproduce those sequences. The student learns from concrete examples of what the teacher tends to produce rather than only from the original human targets or from the teacher’s next-token probabilities.

Artificial Intelligence 09 Sep 2026 9 min read

Use Predictive Entropy to Detect Uncertain Classifications

Use Predictive Entropy to Detect Uncertain Classifications A classifier can return the same predicted label for two inputs while being much less certain about one of them. If an application only keeps the winning label, that difference disappears. Predictive entropy gives you a compact way to preserve it. It summarizes how spread out a classifier’s predicted probability distribution is: concentrated probability produces low entropy, while probability spread across several classes produces higher entropy.

Artificial Intelligence 09 Sep 2026 11 min read

Train with Unlabeled Data Using Mean Teacher

Labeled examples are often the expensive part of an AI system. You may have millions of inputs but only a small subset with trustworthy labels. Training only on the labeled subset ignores information in the rest of the data, while assigning guessed labels too aggressively can teach the model its own mistakes. Mean Teacher is a semi-supervised learning method for this situation. It trains a normal model, called the student, while maintaining a second model, called the teacher, whose parameters are an exponential moving average of the student’s parameters. The student learns from real labels when they exist and is also encouraged to make predictions that agree with the teacher on unlabeled inputs.

Artificial Intelligence 09 Sep 2026 12 min read

Train Discrete Neural Operations with Straight-Through Estimators

Neural networks are usually trained with gradient descent, which depends on small changes in parameters producing informative changes in the loss. A discrete operation can break that assumption. Rounding a value, choosing a binary gate, or selecting a quantized level may be exactly what the forward computation needs, yet its derivative can be zero almost everywhere or undefined at transition points. A straight-through estimator (STE) is a practical way to keep training in that situation. The forward pass uses the discrete operation, while the backward pass substitutes a simpler derivative so that a gradient can flow through it. The important consequence is easy to miss: the backward signal is generally not the true derivative of the discrete forward computation. It is a deliberately chosen surrogate.

Artificial Intelligence 09 Sep 2026 12 min read

Evaluate Sequence Models Beyond Training Length

A sequence model can score well on a random test split and still fail when an input is longer than the sequences it saw during training. This matters for language, symbolic reasoning, event sequences, and other tasks where production inputs do not have one fixed length. The problem is easy to hide. If training and test examples come from the same length distribution, an aggregate metric mostly measures performance on familiar lengths. It does not tell you whether the model learned a rule that extends to longer sequences or a strategy that works only inside the observed range.

Artificial Intelligence 09 Sep 2026 9 min read

Diagnose and Prevent Dying ReLU Units

ReLU is one of the simplest neural-network activation functions: negative inputs become zero and positive inputs pass through unchanged. That simplicity makes optimization efficient, but it creates a failure mode that can quietly waste model capacity. A unit can move into a state where its pre-activation is negative for every relevant training example, so its ReLU output stays zero and the unit stops receiving a useful gradient through that activation.

Artificial Intelligence 09 Sep 2026 10 min read

Detect Representation Collapse in Self-Supervised Learning

Self-supervised learning can train an encoder without manually assigning a class label to every example. But removing labels also removes an obvious force that tells different examples to occupy meaningfully different parts of representation space. A badly designed objective can therefore admit a trivial solution: the encoder maps many or all inputs to essentially the same representation. This failure is called representation collapse. The training loss may even look good, because a model that emits the same vector for two augmented views of every input has achieved perfect agreement without learning useful distinctions.

Artificial Intelligence 09 Sep 2026 11 min read

Control Beam Search Length Bias with Length Normalization

Beam search is a common way to decode sequence models when choosing the most likely token at every step is too shortsighted. It keeps several partial candidates alive, expands them, and repeatedly retains the strongest alternatives. There is a subtle problem: the score used for a sequence usually accumulates one log-probability per generated token. Because token probabilities are at most 1, their log-probabilities are normally non-positive. Extending a sequence therefore tends to make its raw cumulative score smaller. When finished candidates of different lengths compete directly, this can create a preference for outputs that end too early.

Artificial Intelligence 08 Sep 2026 11 min read

Vanishing and Exploding Gradients in Deep Networks

A deep neural network can have enough capacity to solve a task and still fail to learn because useful training signals do not reach all of its layers. Parameters near the output may update normally while earlier layers receive gradients that are almost zero. In the opposite case, gradients can grow so large that one optimizer step destabilizes the model. These are the vanishing-gradient and exploding-gradient problems. They are not simply labels for “training is bad.” They describe what happens to derivatives as backpropagation repeatedly applies the chain rule through many transformations.

Artificial Intelligence 08 Sep 2026 9 min read

Use Test-Time Augmentation Without Hiding Model Errors

A model usually makes one prediction from one representation of an input. That is convenient, but the representation may contain accidental details that should not determine the answer. A product photo can be shifted a few pixels. A scanned digit can be slightly rotated. A crop can place the object closer to one edge than another. Test-time augmentation (TTA) asks the trained model to predict several valid transformations of the same input and then combines those predictions. The technique can make inference less dependent on one particular view, but it also increases compute and can make predictions worse when the transformations change information that matters to the label.

Artificial Intelligence 08 Sep 2026 9 min read

Use Test-Time Augmentation for More Stable Predictions

A classifier can change its prediction because an object moved a few pixels, an image was cropped differently, or another harmless transformation changed the input representation. If those transformations should not change the correct answer, that sensitivity is undesirable. Test-time augmentation (TTA) addresses this problem by running the same trained model on several meaning-preserving versions of an input and combining their predictions. Instead of asking the model for one view of the evidence, TTA asks it to evaluate several valid views.

Artificial Intelligence 08 Sep 2026 10 min read

Train Neural Networks with Curriculum Learning

Most training pipelines treat the dataset as a fixed pool and repeatedly shuffle it. That is a strong default: it is simple, exposes the model to the full data distribution, and avoids assumptions about which examples should come first. But some learning problems have a useful notion of progression. A model may learn basic patterns more reliably before it is asked to handle noisy, ambiguous, or structurally difficult examples. Curriculum learning makes that progression explicit. Instead of changing the model architecture or the loss, it changes which training examples are emphasized at different stages of training. A common curriculum begins with easier examples and gradually introduces harder ones.

Artificial Intelligence 08 Sep 2026 11 min read

Trade Compute for Memory with Activation Checkpointing

A neural network can fit comfortably in accelerator memory for inference and still run out of memory during training. The reason is that training needs more than model weights. Backpropagation also needs intermediate values from the forward pass, and those activations can consume a large share of memory in deep models or with long sequences and large batches. Activation checkpointing reduces that memory pressure by deliberately not keeping every intermediate activation. Instead, training saves selected checkpoints and recomputes missing forward-pass values when the backward pass needs them. The trade is straightforward: keep fewer activations in memory, but perform extra computation.

Artificial Intelligence 08 Sep 2026 9 min read

The Softmax Bottleneck in Language Models

A language model can have a powerful network behind it and still be constrained by the layer that turns its hidden state into next-token probabilities. In the common linear-softmax output layer, that constraint has a precise form: across many contexts, the model can represent only a limited family of log-probability patterns. This limitation is known as the softmax bottleneck. It is not a claim that softmax itself is defective, nor does it mean every modern language model is visibly harmed by it. It is a structural result about a particular output parameterization.

Artificial Intelligence 08 Sep 2026 9 min read

Stabilize Neural Network Weights with Exponential Moving Averages

A neural network’s parameters rarely move smoothly toward their final values. Mini-batch training produces noisy updates: one batch may push a weight in one direction, while the next pushes it partly back. The final checkpoint therefore represents one point on a noisy training path, not necessarily the most useful point near the end of that path. An exponential moving average (EMA) of model weights keeps a second set of parameters that changes more slowly than the actively trained model. Recent training states contribute more than old ones, but no single update immediately replaces the averaged weights.

Artificial Intelligence 08 Sep 2026 8 min read

Residual Connections in Deep Neural Networks

Making a neural network deeper gives it more transformations to work with, but depth alone does not make optimization easy. A stack of layers must learn useful transformations while gradients travel backward through every stage. As the stack grows, that optimization path can become difficult even when the deeper model has enough capacity to represent a good solution. Residual connections change what a block is asked to learn. Instead of making the block produce an entirely new representation, they let it learn a change to the representation it already received. The original input travels along a shortcut and is added back to the learned branch.