Skip to content

Archive

Model Training

55 articles
Artificial Intelligence 24 Sep 2026 6 min read

Packed Sequences Need Boundary-Aware Attention Masks

Sequence packing reduces padding by placing several variable-length examples into one token buffer. The storage layout may look like one long sequence, but the examples are still semantically independent. A standard causal mask does not preserve that independence by itself. For a decoder-only Transformer, causal masking blocks attention to future positions. It does not normally block attention to earlier positions that belong to another packed example. If boundaries are ignored, tokens in a later example can attend to keys and values from an earlier one. The model then receives context that the data pipeline intended to keep separate.

Artificial Intelligence 24 Sep 2026 4 min read

Activation Checkpointing Trades Saved Activations for Recomputation

Activation checkpointing changes which forward-pass tensors remain resident until backpropagation. Instead of retaining every intermediate activation required by gradient computation, a checkpointed region keeps selected boundary state and reconstructs discarded intermediates when the backward pass reaches that region. The mechanism reduces activation memory at the cost of extra computation. It does not shrink model parameters, optimizer state, or gradients, so its effect on total training memory depends on how much of the footprint comes from activations.

Artificial Intelligence 23 Sep 2026 5 min read

Switch Routing Capacity Turns Expert Imbalance into Token Overflow

A Switch-style sparse layer can send many tokens toward the same expert even when every expert has identical compute capacity. The router makes token-level choices from model-produced scores; it does not inherently produce an even partition. A fixed expert capacity therefore creates a hard boundary between routing preference and the amount of expert computation admitted for a batch. In the Switch Transformer formulation, each token is routed to the expert with the highest router probability. Each expert receives a fixed token capacity derived from the token count, expert count, and a capacity factor. When assignments exceed that capacity, excess tokens overflow instead of enlarging the expert batch without bound.

Artificial Intelligence 23 Sep 2026 5 min read

Label Smoothing Redistributes Target Probability Before Cross-Entropy

A classifier trained with one-hot targets is asked to place all target probability on a single class. Cross-entropy does not require that target representation. Label smoothing changes the target distribution before the loss is evaluated, so the model receives a different gradient even when its logits and predicted probabilities are unchanged. This distinction matters because label smoothing is not a decoding rule and does not alter inference by itself. It changes the training objective. The resulting model parameters can differ because the optimizer follows gradients computed against softened targets.

Artificial Intelligence 23 Sep 2026 5 min read

Global Gradient Clipping Caps Norm Before the Optimizer Step

A single training step can produce gradients whose combined magnitude is far larger than nearby steps. Global norm clipping changes that gradient set before the optimizer consumes it. When the measured norm exceeds a configured threshold, every selected gradient is multiplied by the same scale factor. The mechanism is simple, but its boundary matters. Clipping controls the norm of the gradients supplied to the optimizer. It does not directly impose the same bound on the eventual parameter update, especially when the optimizer keeps momentum or adaptive state.

Artificial Intelligence 17 Sep 2026 6 min read

Limit Parameter Drift with Elastic Weight Consolidation

Fine-tuning a neural network on a new task can move parameters away from values that supported an earlier task. The new objective has no inherent reason to preserve those earlier behaviors when the earlier data is absent from the update. Elastic weight consolidation, commonly abbreviated EWC, adds a parameter-space constraint intended to reduce that drift. The constraint is selective rather than uniform. Parameters estimated to be more consequential for the earlier task receive a larger penalty for moving, while parameters assigned lower importance can change more freely. That distinction is the central mechanism; EWC is not simply weight decay around zero.

Artificial Intelligence 17 Sep 2026 7 min read

Balance Token Routing in Sparse Mixture-of-Experts Models

A sparse mixture-of-experts layer can contain many expert networks while evaluating only a small subset for each token. The router makes that sparsity possible: it assigns scores to experts, selects a limited set, and sends each token through the selected computation paths. That selection is not only an optimization detail. If many tokens concentrate on a few experts, some devices can receive much more work than others, capacity limits can discard or redirect assignments, and experts that receive little traffic get fewer task gradients. Router balance therefore affects both computation and the function represented by the model.

Artificial Intelligence 16 Sep 2026 6 min read

Trade Activation Memory for Recomputation

Backpropagation needs intermediate values from the forward computation to form gradients. Retaining every required activation can consume substantial accelerator memory, especially as sequence length, batch size, hidden width, or network depth grows. Activation checkpointing changes that storage decision. Selected forward regions retain only chosen boundary tensors, then reproduce omitted intermediates when the backward pass reaches those regions. Peak activation memory can fall, but some forward computation is executed again. The useful engineering question is not simply whether checkpointing saves memory. The placement of recomputation boundaries determines which tensors disappear, how much extra compute appears, and whether replayed operations reproduce a valid backward computation.

Artificial Intelligence 16 Sep 2026 6 min read

Measure Gradient Noise Before Scaling Batch Size

Increasing a training batch reduces variation in the minibatch gradient, but the reduction does not continue to buy proportional progress indefinitely. Once a batch is large enough that its gradient estimate is already dominated by the underlying gradient signal, processing more examples before the next parameter update yields diminishing algorithmic returns. Gradient noise scale gives this transition a measurable form. It compares stochastic variation in per-example gradients with the magnitude of the mean gradient. The quantity is not a universal batch-size setting, and its exact estimator depends on assumptions about sampling and gradient aggregation. It is useful as a diagnostic for how much additional batch parallelism the current optimization state can absorb.

Artificial Intelligence 16 Sep 2026 6 min read

Isolate Packed Sequences During Transformer Training

Padding can consume a large share of a training batch when sequence lengths vary. Sequence packing replaces some of that padding with tokens from additional examples, placing multiple independent samples inside one fixed-length token block. The arithmetic is attractive: more of each block carries data that contributes to the training objective. The packed tensor, however, no longer describes one continuous sequence. If the model treats it that way, tokens from a later example can attend to tokens from an earlier one. The optimizer then sees dependencies that were absent from the original dataset. Packing is therefore not only a batching optimization. It changes the structure presented to the attention mechanism unless example boundaries are represented explicitly.

Artificial Intelligence 16 Sep 2026 6 min read

Bucket Sequence Lengths to Reduce Padding Waste

A padded batch is shaped by its longest sequence, not its average sequence. If one batch contains token counts of 120, 124, 131, and 900, every sequence may be represented at length 900. Most positions in the first three rows then carry padding rather than input tokens. Length bucketing changes batch composition instead of changing the model. Examples with similar token counts are placed near each other before batches are formed. The maximum length inside each batch falls closer to the lengths of its members, reducing the number of padded positions processed by operations that still use the rectangular batch shape.

Artificial Intelligence 15 Sep 2026 5 min read

Isolate Attention Across Packed Training Sequences

Padding can consume a large share of a training batch when sequence lengths vary. Sequence packing reduces that waste by placing multiple short samples into one token block. The arithmetic is attractive: more non-padding tokens fit into the same fixed-length tensor. Packing also changes the structure seen by attention. A standard causal mask only prevents a token from attending to later positions. It does not know that two adjacent spans came from separate samples. Without an additional boundary constraint, a token in the second span can attend to tokens from the first span.

Artificial Intelligence 15 Sep 2026 6 min read

Clip Gradient Norms With Clear Scope

Gradient norm clipping changes an optimizer update only when the measured gradient norm exceeds a chosen threshold. The operation sounds local, but its behavior depends on a broader implementation choice: which gradients participate in the norm. Two training loops can use the same threshold and optimizer yet produce different updates because they clip different parameter groups or clip at different points in the update cycle. That makes clipping scope part of the optimization definition, not just a guard against unusually large gradients.

Artificial Intelligence 15 Sep 2026 6 min read

Account for Label Smoothing in Classifier Confidence

A classifier trained with one-hot targets is rewarded for moving probability mass toward the labeled class. Cross-entropy keeps decreasing as the model assigns that class a probability closer to one, even after the predicted class is already correct. Label smoothing changes this pressure by assigning a small amount of target mass to the other classes. That change is easy to treat as a minor detail in the loss function. It is not minor when an application consumes the model’s probability values. The smoothed target changes the optimum encouraged by the training objective, so confidence scores from a smoothed model should not be interpreted as if they came from the same objective as ordinary one-hot training.

Artificial Intelligence 14 Sep 2026 7 min read

Trade Activation Memory for Recomputation with Checkpointing

Backpropagation needs intermediate values from the forward pass to compute parameter and input gradients. Keeping every required activation alive until its gradient is calculated can consume substantial device memory, especially as model depth, batch size, or sequence length grows. Activation checkpointing changes which intermediates are retained. Selected boundary tensors remain available, while activations inside a checkpointed region are discarded after the forward pass and produced again when the backward pass reaches that region. The model computes the same conceptual function, but the execution schedule exchanges additional computation for lower activation storage.

Artificial Intelligence 14 Sep 2026 5 min read

Prevent Cross-Example Attention in Packed Sequences

Short training examples can waste much of a fixed-length transformer batch on padding. Sequence packing reduces that waste by placing several examples into one token buffer, but concatenation alone changes the computation. A causal mask prevents a token from attending to future positions; it does not prevent that token from attending to an earlier, unrelated example. The distinction matters whenever packed examples are intended to remain independent. The token buffer may be contiguous for storage and compute while attention, position handling, and loss accounting still need explicit example boundaries.

Artificial Intelligence 14 Sep 2026 7 min read

Control Target Certainty with Label Smoothing

A classifier trained with ordinary cross-entropy often receives a one-hot target: probability mass 1 on the labeled class and 0 on every other class. That target keeps rewarding movement toward a more extreme prediction even after the correct class already has the highest score. Label smoothing changes the target distribution rather than the model architecture. A small amount of target mass is moved away from the labeled class and assigned to other classes. Cross-entropy then optimizes against this softened distribution, so the gradient no longer treats absolute certainty on the labeled class as the target state.

Artificial Intelligence 14 Sep 2026 6 min read

Accumulate Gradients Across Microbatches

A training batch can exceed accelerator memory even when model parameters and optimizer state fit comfortably. Activations from the forward pass often account for a large part of the remaining footprint, and their memory cost grows with the number of examples processed together. Gradient accumulation splits a larger logical batch into smaller microbatches. Each microbatch runs its own forward and backward pass, but the optimizer waits until several backward passes have contributed to the parameter gradients. This reduces the activation memory required for any single pass without requiring an optimizer update after every microbatch.

Artificial Intelligence 11 Sep 2026 9 min read

Use Token Dropout to Train Robust Sequence Models

Use Token Dropout to Train Robust Sequence Models A sequence model can become too dependent on a few easy input clues. Remove one field, truncate a message, or corrupt a token at inference time, and a prediction that looked reliable on clean validation data may change sharply. Token dropout is a simple training-time corruption technique: randomly hide some input tokens while keeping the learning target unchanged. The model is forced to solve some training examples without every usual clue. Used carefully, this can reduce brittle dependence on individual tokens. Used carelessly, it can destroy information the task genuinely requires.

Artificial Intelligence 11 Sep 2026 10 min read

Train Through Discrete Decisions with the Straight-Through Estimator

Train Through Discrete Decisions with the Straight-Through Estimator Neural networks are usually trained by following gradients through a chain of differentiable operations. A hard discrete choice breaks that chain. Rounding a value, selecting a binary gate, or quantizing an activation can make the forward computation useful while leaving ordinary backpropagation with a zero or undefined derivative at the decision. The straight-through estimator (STE) is a practical workaround. It keeps the discrete operation in the forward pass but substitutes a simpler derivative during the backward pass. That makes optimization possible, at the cost of using a gradient that is not the true derivative of the forward computation.

Artificial Intelligence 11 Sep 2026 9 min read

Score Candidates with Energy-Based Models

Score Candidates with Energy-Based Models Many AI systems need to decide which candidate fits an input: which reply matches a conversation, which label fits an image, or which configuration is plausible. A common design makes the model output a probability directly. Energy-based models take a more general route: they assign each input-candidate pair a scalar energy, with lower values representing greater compatibility. That simple change is useful because the model can focus on relative preference without requiring every architecture to produce a normalized probability during scoring. It also introduces real engineering challenges. Training needs informative alternatives, probability normalization can be expensive, and inference may require searching a large candidate space.

Artificial Intelligence 11 Sep 2026 11 min read

Reweight Long-Tailed Classification with Effective Sample Counts

A classifier trained on a long-tailed dataset can see thousands of examples from common classes and only a handful from rare ones. Ordinary empirical risk minimization gives the common classes more influence simply because they appear more often. A tempting fix is to weight each class by the inverse of its example count, but that can make a tiny class disproportionately influential, including any mislabeled examples it contains. Class-balanced loss based on the effective number of samples provides a smoother way to derive class weights. Instead of treating every additional example as equally informative, it models diminishing returns within a class and weights classes according to an adjusted, or effective, sample count.

Artificial Intelligence 11 Sep 2026 9 min read

Focus Classifier Training with Focal Loss

Focus Classifier Training with Focal Loss A classifier can spend much of its training signal on examples it already handles correctly. This becomes especially troublesome when easy examples vastly outnumber difficult ones. A detector, for instance, may encounter many obvious background locations for every location containing an object. Focal loss changes the contribution of each example according to the model’s confidence in the correct class. Easy, high-confidence examples receive less weight. Harder examples retain more of their cross-entropy loss. The mechanism is small, but using it well requires understanding what it changes and what it doesn’t.

Artificial Intelligence 11 Sep 2026 9 min read

Control Neural Network Weight Scale with Spectral Normalization

Control Neural Network Weight Scale with Spectral Normalization A neural network layer can amplify a small change in its input into a much larger change in its output. Large amplification isn’t automatically a defect, but it can make some models harder to control during training, especially when one network is reacting to another as in a generative adversarial network. Spectral normalization puts a direct constraint on that amplification for a linear transformation. It rescales a weight matrix using its largest singular value, called the spectral norm. The result is a simple mechanism with a precise local meaning: under the Euclidean norm, the normalized linear map cannot stretch a vector by more than the chosen scale.