Skip to content

Archive

Neural Networks

96 articles
Artificial Intelligence 24 Sep 2026 4 min read

Activation Checkpointing Trades Saved Activations for Recomputation

Activation checkpointing changes which forward-pass tensors remain resident until backpropagation. Instead of retaining every intermediate activation required by gradient computation, a checkpointed region keeps selected boundary state and reconstructs discarded intermediates when the backward pass reaches that region. The mechanism reduces activation memory at the cost of extra computation. It does not shrink model parameters, optimizer state, or gradients, so its effect on total training memory depends on how much of the footprint comes from activations.

Artificial Intelligence 23 Sep 2026 5 min read

Global Gradient Clipping Caps Norm Before the Optimizer Step

A single training step can produce gradients whose combined magnitude is far larger than nearby steps. Global norm clipping changes that gradient set before the optimizer consumes it. When the measured norm exceeds a configured threshold, every selected gradient is multiplied by the same scale factor. The mechanism is simple, but its boundary matters. Clipping controls the norm of the gradients supplied to the optimizer. It does not directly impose the same bound on the eventual parameter update, especially when the optimizer keeps momentum or adaptive state.

Artificial Intelligence 23 Sep 2026 5 min read

Activation Outliers Can Set the Scale for an Entire Quantization Group

A quantizer with a fixed integer width has only a finite set of representable codes. When many activations share one scale, a single value with much larger magnitude can force that scale to cover a wider real-valued range. The remaining values then occupy fewer useful code intervals around the region where they are concentrated. This behavior is not a generic statement that quantization fails in the presence of large numbers. It follows from a specific coupling: values inside the same quantization group share parameters that map real numbers to integer codes.

Artificial Intelligence 18 Sep 2026 5 min read

Recompute BatchNorm Statistics After Weight Averaging

Averaging two neural-network checkpoints can produce a useful parameter vector, yet leave BatchNorm running statistics tied to a different network. The weights define one set of activations; the stored running means and variances may describe activations produced by earlier weights. Inference then combines state from two different points in parameter space. This mismatch is easy to miss because BatchNorm running statistics are buffers rather than trainable parameters in common implementations. A parameter-averaging routine can handle every weight correctly and still produce an internally inconsistent inference state.

Artificial Intelligence 17 Sep 2026 6 min read

Limit Parameter Drift with Elastic Weight Consolidation

Fine-tuning a neural network on a new task can move parameters away from values that supported an earlier task. The new objective has no inherent reason to preserve those earlier behaviors when the earlier data is absent from the update. Elastic weight consolidation, commonly abbreviated EWC, adds a parameter-space constraint intended to reduce that drift. The constraint is selective rather than uniform. Parameters estimated to be more consequential for the earlier task receive a larger penalty for moving, while parameters assigned lower importance can change more freely. That distinction is the central mechanism; EWC is not simply weight decay around zero.

Artificial Intelligence 17 Sep 2026 6 min read

Bound Gradient Updates with Global Norm Clipping

A single optimization step can contain gradients whose combined magnitude is far larger than the surrounding steps. If those gradients are passed directly to an optimizer, the resulting parameter update can move the model into a very different region of parameter space. Global norm clipping places a bound on the gradient magnitude before the optimizer consumes it. The mechanism is simple, but its behavior is easy to misread. It does not cap every gradient element independently, and it does not guarantee a fixed parameter-update norm for adaptive optimizers. It rescales the collected gradient vector when a chosen norm crosses a threshold.

Artificial Intelligence 16 Sep 2026 6 min read

Trade Activation Memory for Recomputation

Backpropagation needs intermediate values from the forward computation to form gradients. Retaining every required activation can consume substantial accelerator memory, especially as sequence length, batch size, hidden width, or network depth grows. Activation checkpointing changes that storage decision. Selected forward regions retain only chosen boundary tensors, then reproduce omitted intermediates when the backward pass reaches those regions. Peak activation memory can fall, but some forward computation is executed again. The useful engineering question is not simply whether checkpointing saves memory. The placement of recomputation boundaries determines which tensors disappear, how much extra compute appears, and whether replayed operations reproduce a valid backward computation.

Artificial Intelligence 16 Sep 2026 6 min read

Measure Gradient Noise Before Scaling Batch Size

Increasing a training batch reduces variation in the minibatch gradient, but the reduction does not continue to buy proportional progress indefinitely. Once a batch is large enough that its gradient estimate is already dominated by the underlying gradient signal, processing more examples before the next parameter update yields diminishing algorithmic returns. Gradient noise scale gives this transition a measurable form. It compares stochastic variation in per-example gradients with the magnitude of the mean gradient. The quantity is not a universal batch-size setting, and its exact estimator depends on assumptions about sampling and gradient aggregation. It is useful as a diagnostic for how much additional batch parallelism the current optimization state can absorb.

Artificial Intelligence 15 Sep 2026 6 min read

Clip Gradient Norms With Clear Scope

Gradient norm clipping changes an optimizer update only when the measured gradient norm exceeds a chosen threshold. The operation sounds local, but its behavior depends on a broader implementation choice: which gradients participate in the norm. Two training loops can use the same threshold and optimizer yet produce different updates because they clip different parameter groups or clip at different points in the update cycle. That makes clipping scope part of the optimization definition, not just a guard against unusually large gradients.

Artificial Intelligence 15 Sep 2026 6 min read

Account for Label Smoothing in Classifier Confidence

A classifier trained with one-hot targets is rewarded for moving probability mass toward the labeled class. Cross-entropy keeps decreasing as the model assigns that class a probability closer to one, even after the predicted class is already correct. Label smoothing changes this pressure by assigning a small amount of target mass to the other classes. That change is easy to treat as a minor detail in the loss function. It is not minor when an application consumes the model’s probability values. The smoothed target changes the optimum encouraged by the training objective, so confidence scores from a smoothed model should not be interpreted as if they came from the same objective as ordinary one-hot training.

Artificial Intelligence 14 Sep 2026 7 min read

Trade Activation Memory for Recomputation with Checkpointing

Backpropagation needs intermediate values from the forward pass to compute parameter and input gradients. Keeping every required activation alive until its gradient is calculated can consume substantial device memory, especially as model depth, batch size, or sequence length grows. Activation checkpointing changes which intermediates are retained. Selected boundary tensors remain available, while activations inside a checkpointed region are discarded after the forward pass and produced again when the backward pass reaches that region. The model computes the same conceptual function, but the execution schedule exchanges additional computation for lower activation storage.

Artificial Intelligence 14 Sep 2026 6 min read

Track Model Weights with an Exponential Moving Average

Optimizer updates can move model parameters back and forth even when the broader trajectory changes more gradually. An exponential moving average, or EMA, keeps a second parameter state that follows those updates with smoothing. The trainable model still receives ordinary optimizer updates; the EMA state is a derived copy used separately, often for evaluation or export. The mechanism is compact, but its behavior depends on decay, update frequency, initialization, and which state is actually saved or evaluated.

Artificial Intelligence 14 Sep 2026 5 min read

Tie Input Embeddings to the Output Projection

A language model can contain two large matrices indexed by the same vocabulary: one maps token IDs into embedding vectors, while another maps hidden states into vocabulary logits. Weight tying makes those roles share parameters instead of maintaining two independent matrices. The change is compact in code, but it affects parameter counting, gradient flow, dimensional constraints, checkpoint handling, and any component that assumes the input and output weights are separate.

Artificial Intelligence 14 Sep 2026 5 min read

Normalize Hidden States with RMSNorm

A hidden-state vector can grow or shrink in magnitude as it passes through a neural network. RMSNorm controls that scale by dividing the vector by its root mean square magnitude, then applying a trainable gain. Unlike LayerNorm, it does not subtract the vector mean before rescaling. That missing centering operation is the defining distinction. RMSNorm constrains scale while leaving a uniform shift across coordinates present in the normalized representation. RMSNorm uses the second raw moment For a hidden vector x with d coordinates, its root mean square is:

Artificial Intelligence 14 Sep 2026 7 min read

Control Target Certainty with Label Smoothing

A classifier trained with ordinary cross-entropy often receives a one-hot target: probability mass 1 on the labeled class and 0 on every other class. That target keeps rewarding movement toward a more extreme prediction even after the correct class already has the highest score. Label smoothing changes the target distribution rather than the model architecture. A small amount of target mass is moved away from the labeled class and assigned to other classes. Cross-entropy then optimizes against this softened distribution, so the gradient no longer treats absolute certainty on the labeled class as the target state.

Artificial Intelligence 14 Sep 2026 5 min read

Control Classifier Logits with Cosine Normalization

A linear classification head mixes two signals in each logit: the angle between a feature vector and a class weight vector, and the magnitudes of both vectors. Cosine normalization removes the magnitude terms, so class scores depend on directional alignment instead. That change is small in code but substantial in interpretation. Feature norm no longer increases every class comparison merely by growing, class-weight norm no longer acts as an implicit class-specific scale, and the overall sharpness of the softmax must be supplied separately.

Artificial Intelligence 14 Sep 2026 6 min read

Compare RMSNorm and Layer Normalization

Normalization layers can look interchangeable when their outputs have similar shapes, but their invariances are not the same. RMSNorm rescales an activation vector using its root mean square without first subtracting the vector mean. Layer normalization centers the vector and then rescales it using its variance. That missing centering operation is the central distinction. It changes which transformations of an activation vector disappear under normalization and which remain visible to the rest of the network.

Artificial Intelligence 14 Sep 2026 6 min read

Accumulate Gradients Across Microbatches

A training batch can exceed accelerator memory even when model parameters and optimizer state fit comfortably. Activations from the forward pass often account for a large part of the remaining footprint, and their memory cost grows with the number of examples processed together. Gradient accumulation splits a larger logical batch into smaller microbatches. Each microbatch runs its own forward and backward pass, but the optimizer waits until several backward passes have contributed to the parameter gradients. This reduces the activation memory required for any single pass without requiring an optimizer update after every microbatch.

Artificial Intelligence 13 Sep 2026 6 min read

Trade Activation Memory for Recomputation with Gradient Checkpointing

Training a deep neural network requires more memory than its parameters alone suggest. Backpropagation needs intermediate values from the forward pass, and retaining those activations across many layers can consume a large share of accelerator memory. Gradient checkpointing changes that storage policy. Instead of retaining every intermediate activation until its gradient is computed, training keeps selected boundary tensors and reconstructs omitted intermediates by running parts of the forward computation again during the backward pass. The model function need not change, but the execution schedule does.

Artificial Intelligence 13 Sep 2026 6 min read

Regularize Classifier Targets with Label Smoothing

A classifier trained with one-hot targets is rewarded for pushing the target class probability toward one and every other class probability toward zero. Cross-entropy keeps applying pressure in that direction even after the predicted class is already correct. Label smoothing changes that pressure by replacing the exact one-hot target with a distribution that reserves some mass for other classes. That small change affects more than the target tensor. It changes the gradient on every output logit, limits the incentive for extreme class separation, and alters how predicted probabilities should be interpreted.

Artificial Intelligence 13 Sep 2026 5 min read

Focus Classification Loss with Focal Modulation

Cross-entropy gives every classified example a loss determined by the probability assigned to its target class. When a training batch contains many examples the model already classifies with high confidence, their individual losses may be small yet their aggregate contribution can still occupy a substantial part of the objective. Focal loss changes that balance with a confidence-dependent multiplier. The mechanism is not a new classifier head or sampling strategy. It modifies the loss so that examples with high target-class probability are attenuated more strongly than examples with low target-class probability.

Artificial Intelligence 13 Sep 2026 6 min read

Detect Out-of-Distribution Inputs with Classifier Energy Scores

A classifier can assign high softmax confidence to an input that does not resemble the data used to fit its parameters. Softmax normalizes scores across the available classes; it does not add a separate class for unfamiliar inputs. As a result, a large maximum probability is not evidence that an input belongs to the expected data distribution. Energy-based out-of-distribution detection uses the full logit vector to produce a scalar score before a deployment policy decides whether an input looks familiar enough to accept. The score is simple to compute for an existing classifier, but its interpretation depends on the model, temperature, data regime, and threshold calibration.

Artificial Intelligence 13 Sep 2026 6 min read

Compare RMSNorm and LayerNorm in Transformers

LayerNorm and RMSNorm can occupy the same structural position in a transformer while applying different operations to the residual stream. LayerNorm subtracts the feature mean before scaling by a measure of spread. RMSNorm skips the centering operation and scales directly from the root mean square of the features. That small algebraic difference changes which transformations of an activation vector are removed by normalization. It also means that replacing one operation with the other is not, in general, a function-preserving edit to an existing model.

Artificial Intelligence 13 Sep 2026 6 min read

Clip Gradient Norms Before Optimizer Updates

A training update can become dominated by a gradient whose magnitude is far larger than the range seen in nearby iterations. Global norm clipping places a bound on that update signal before the optimizer consumes it. The operation is simple, but its behavior depends on what is included in the norm, where clipping occurs, and how it interacts with gradient accumulation and mixed-precision scaling. Norm clipping does not repair the source of unstable gradients. It changes the vector passed to the optimizer when its norm exceeds a chosen threshold.