Skip to content

Archive

Optimization

24 articles
Artificial Intelligence 24 Sep 2026 4 min read

Global Gradient Norm Clipping Rescales Parameter Gradients with One Shared Factor

Global gradient norm clipping treats the current parameter gradients as one aggregate vector. If its norm exceeds a threshold, every participating gradient is multiplied by the same coefficient. The operation changes update magnitude while preserving the direction of the aggregate gradient vector, apart from finite-precision effects. One norm controls the whole gradient set Let the parameter gradients be g_1, g_2, ... g_m. For a p-norm, global clipping first forms the equivalent norm of their concatenation:

Artificial Intelligence 23 Sep 2026 5 min read

Label Smoothing Redistributes Target Probability Before Cross-Entropy

A classifier trained with one-hot targets is asked to place all target probability on a single class. Cross-entropy does not require that target representation. Label smoothing changes the target distribution before the loss is evaluated, so the model receives a different gradient even when its logits and predicted probabilities are unchanged. This distinction matters because label smoothing is not a decoding rule and does not alter inference by itself. It changes the training objective. The resulting model parameters can differ because the optimizer follows gradients computed against softened targets.

Artificial Intelligence 23 Sep 2026 5 min read

Global Gradient Clipping Caps Norm Before the Optimizer Step

A single training step can produce gradients whose combined magnitude is far larger than nearby steps. Global norm clipping changes that gradient set before the optimizer consumes it. When the measured norm exceeds a configured threshold, every selected gradient is multiplied by the same scale factor. The mechanism is simple, but its boundary matters. Clipping controls the norm of the gradients supplied to the optimizer. It does not directly impose the same bound on the eventual parameter update, especially when the optimizer keeps momentum or adaptive state.

Artificial Intelligence 18 Sep 2026 5 min read

Recompute BatchNorm Statistics After Weight Averaging

Averaging two neural-network checkpoints can produce a useful parameter vector, yet leave BatchNorm running statistics tied to a different network. The weights define one set of activations; the stored running means and variances may describe activations produced by earlier weights. Inference then combines state from two different points in parameter space. This mismatch is easy to miss because BatchNorm running statistics are buffers rather than trainable parameters in common implementations. A parameter-averaging routine can handle every weight correctly and still produce an internally inconsistent inference state.

Artificial Intelligence 17 Sep 2026 6 min read

Limit Parameter Drift with Elastic Weight Consolidation

Fine-tuning a neural network on a new task can move parameters away from values that supported an earlier task. The new objective has no inherent reason to preserve those earlier behaviors when the earlier data is absent from the update. Elastic weight consolidation, commonly abbreviated EWC, adds a parameter-space constraint intended to reduce that drift. The constraint is selective rather than uniform. Parameters estimated to be more consequential for the earlier task receive a larger penalty for moving, while parameters assigned lower importance can change more freely. That distinction is the central mechanism; EWC is not simply weight decay around zero.

Artificial Intelligence 17 Sep 2026 6 min read

Bound Gradient Updates with Global Norm Clipping

A single optimization step can contain gradients whose combined magnitude is far larger than the surrounding steps. If those gradients are passed directly to an optimizer, the resulting parameter update can move the model into a very different region of parameter space. Global norm clipping places a bound on the gradient magnitude before the optimizer consumes it. The mechanism is simple, but its behavior is easy to misread. It does not cap every gradient element independently, and it does not guarantee a fixed parameter-update norm for adaptive optimizers. It rescales the collected gradient vector when a chosen norm crosses a threshold.

Artificial Intelligence 16 Sep 2026 6 min read

Measure Gradient Noise Before Scaling Batch Size

Increasing a training batch reduces variation in the minibatch gradient, but the reduction does not continue to buy proportional progress indefinitely. Once a batch is large enough that its gradient estimate is already dominated by the underlying gradient signal, processing more examples before the next parameter update yields diminishing algorithmic returns. Gradient noise scale gives this transition a measurable form. It compares stochastic variation in per-example gradients with the magnitude of the mean gradient. The quantity is not a universal batch-size setting, and its exact estimator depends on assumptions about sampling and gradient aggregation. It is useful as a diagnostic for how much additional batch parallelism the current optimization state can absorb.

Artificial Intelligence 15 Sep 2026 6 min read

Clip Gradient Norms With Clear Scope

Gradient norm clipping changes an optimizer update only when the measured gradient norm exceeds a chosen threshold. The operation sounds local, but its behavior depends on a broader implementation choice: which gradients participate in the norm. Two training loops can use the same threshold and optimizer yet produce different updates because they clip different parameter groups or clip at different points in the update cycle. That makes clipping scope part of the optimization definition, not just a guard against unusually large gradients.

Artificial Intelligence 14 Sep 2026 6 min read

Track Model Weights with an Exponential Moving Average

Optimizer updates can move model parameters back and forth even when the broader trajectory changes more gradually. An exponential moving average, or EMA, keeps a second parameter state that follows those updates with smoothing. The trainable model still receives ordinary optimizer updates; the EMA state is a derived copy used separately, often for evaluation or export. The mechanism is compact, but its behavior depends on decay, update frequency, initialization, and which state is actually saved or evaluated.

Artificial Intelligence 14 Sep 2026 6 min read

Accumulate Gradients Across Microbatches

A training batch can exceed accelerator memory even when model parameters and optimizer state fit comfortably. Activations from the forward pass often account for a large part of the remaining footprint, and their memory cost grows with the number of examples processed together. Gradient accumulation splits a larger logical batch into smaller microbatches. Each microbatch runs its own forward and backward pass, but the optimizer waits until several backward passes have contributed to the parameter gradients. This reduces the activation memory required for any single pass without requiring an optimizer update after every microbatch.

Artificial Intelligence 13 Sep 2026 6 min read

Clip Gradient Norms Before Optimizer Updates

A training update can become dominated by a gradient whose magnitude is far larger than the range seen in nearby iterations. Global norm clipping places a bound on that update signal before the optimizer consumes it. The operation is simple, but its behavior depends on what is included in the norm, where clipping occurs, and how it interacts with gradient accumulation and mixed-precision scaling. Norm clipping does not repair the source of unstable gradients. It changes the vector passed to the optimizer when its norm exceeds a chosen threshold.

Artificial Intelligence 11 Sep 2026 10 min read

Understand Wide Neural Networks with the Neural Tangent Kernel

Understand Wide Neural Networks with the Neural Tangent Kernel A neural network may contain millions of parameters, yet a useful theoretical view asks a smaller question: when one training example changes the parameters, how does that update affect the prediction for another example? The neural tangent kernel (NTK) answers that question through gradients. It measures how similarly two inputs respond to an infinitesimal parameter update. In a particular infinite-width regime, this kernel becomes effectively fixed during training, turning a nonlinear parameter-optimization problem into a much simpler kernel process in function space.

Artificial Intelligence 11 Sep 2026 10 min read

Train Through Discrete Decisions with the Straight-Through Estimator

Train Through Discrete Decisions with the Straight-Through Estimator Neural networks are usually trained by following gradients through a chain of differentiable operations. A hard discrete choice breaks that chain. Rounding a value, selecting a binary gate, or quantizing an activation can make the forward computation useful while leaving ordinary backpropagation with a zero or undefined derivative at the decision. The straight-through estimator (STE) is a practical workaround. It keeps the discrete operation in the forward pass but substitutes a simpler derivative during the backward pass. That makes optimization possible, at the cost of using a gradient that is not the true derivative of the forward computation.

Artificial Intelligence 11 Sep 2026 11 min read

Clip Gradients Relative to Parameter Scale with AGC

Clip Gradients Relative to Parameter Scale with AGC A gradient can be large in absolute terms without being large for the parameter it updates. A gradient norm of 0.1 is modest next to a parameter norm of 10, but enormous next to a parameter norm of 0.001. Ordinary gradient clipping doesn’t see that distinction: it compares gradients with a fixed threshold. Adaptive gradient clipping (AGC) uses a different reference point. It compares a gradient’s norm with the norm of the parameter unit that gradient will update. If the gradient is too large relative to the parameter, AGC rescales it before the optimizer step.

Artificial Intelligence 09 Sep 2026 12 min read

Train Discrete Neural Operations with Straight-Through Estimators

Neural networks are usually trained with gradient descent, which depends on small changes in parameters producing informative changes in the loss. A discrete operation can break that assumption. Rounding a value, choosing a binary gate, or selecting a quantized level may be exactly what the forward computation needs, yet its derivative can be zero almost everywhere or undefined at transition points. A straight-through estimator (STE) is a practical way to keep training in that situation. The forward pass uses the discrete operation, while the backward pass substitutes a simpler derivative so that a gradient can flow through it. The important consequence is easy to miss: the backward signal is generally not the true derivative of the discrete forward computation. It is a deliberately chosen surrogate.

Artificial Intelligence 07 Sep 2026 12 min read

Xavier and He Initialization for Neural Networks

A deep neural network can fail before learning has had a fair chance. If its initial weights make activations or gradients shrink layer after layer, useful signals can become tiny. If those quantities grow too much, training can become unstable. The optimizer may receive a problem that is unnecessarily difficult even though the architecture and data are otherwise reasonable. Weight initialization tries to start the network in a numerically useful regime. Two common schemes are Xavier initialization, also called Glorot initialization, and He initialization, also called Kaiming initialization. Both choose the scale of random weights from the size of a layer, but they make different assumptions about how signals pass through the activation function.

Artificial Intelligence 07 Sep 2026 9 min read

Weight Decay in Adam and AdamW

A training configuration can contain a parameter named weight_decay without making it obvious what operation the optimizer actually performs. That ambiguity matters most with adaptive optimizers such as Adam: adding an L2 penalty to the loss and directly decaying parameters are not generally the same update. The distinction is easy to miss because the two procedures are closely related under ordinary stochastic gradient descent (SGD). Once an optimizer rescales different coordinates using gradient history, however, the equivalence breaks.

Artificial Intelligence 07 Sep 2026 10 min read

Train Neural Networks with Sharpness-Aware Minimization

A neural network can reach low training loss at parameter values where a small change to the weights makes the loss rise sharply. Standard optimization does not directly discourage this behavior: it mainly asks whether the loss is low at the current parameters. Sharpness-Aware Minimization (SAM) changes the training objective. Instead of optimizing only the loss at the current weights, it approximately optimizes the worst loss in a small neighborhood around them. The practical idea is simple: find a nearby parameter perturbation that makes the current mini-batch harder, then update the original model using the gradient measured at those perturbed parameters.

Artificial Intelligence 07 Sep 2026 10 min read

Resolve Conflicting Gradients in Multi-Task Learning

Training one neural network to solve several tasks can reduce duplicated computation and let related tasks share useful representations. It also creates a problem that single-task training does not have: two losses can ask the same shared parameter to move in opposing directions during the same update. Simply adding the losses does not make that disagreement disappear. Their gradients are added too, so one task can partially cancel another or dominate the shared update. Gradient surgery is a family of techniques that changes task gradients before combining them. A well-known example is projected conflicting gradients, commonly called PCGrad, which removes a conflicting component of one task’s gradient relative to another.

Artificial Intelligence 07 Sep 2026 12 min read

Gradient Noise Scale and Batch Size

Increasing a neural network’s batch size can make more accelerators useful, but the benefit does not grow indefinitely. At some point, processing more examples before each update gives a cleaner estimate of nearly the same gradient direction while consuming additional examples and compute. Gradient noise scale provides a useful mental model for this transition. It compares the variation in per-example gradients with the strength of their average. When gradient estimates are noisy relative to their mean, averaging more examples can remove meaningful noise. When they are already stable, a larger batch has less statistical work left to do.

Artificial Intelligence 07 Sep 2026 9 min read

Gradient Centralization in Neural Network Training

Neural network optimizers normally consume the gradients produced by backpropagation directly. Gradient centralization inserts one small transformation between those two steps: for selected weight tensors, it subtracts the mean of each gradient vector before the optimizer uses it. That operation is easy to implement, but its effect is easy to misunderstand. It is not gradient clipping, because it does not cap large values. It is not normalization, because it does not divide by a norm or standard deviation. It changes the direction of the update by removing one particular component.

Artificial Intelligence 07 Sep 2026 10 min read

Average Neural Network Weights with Stochastic Weight Averaging

A neural network rarely finishes training at the only useful point in parameter space. Late in training, stochastic gradient descent (SGD) can visit several nearby parameter settings that all perform reasonably well, while the final checkpoint represents only one of them. Stochastic weight averaging (SWA) turns that observation into a simple training technique: collect model parameters from multiple points late in an SGD trajectory and compute their arithmetic mean. The result is one model with averaged weights, so inference does not require running an ensemble of all collected checkpoints.

Artificial Intelligence 06 Sep 2026 10 min read

Reduce Transformer Inference with Early Exits

A transformer classifier normally spends the same number of layers on every input. A straightforward support ticket and an ambiguous one both travel through the entire network, even when an intermediate representation may already contain enough information for the easy case. Early exiting changes that fixed-compute rule. It adds prediction points inside the model and lets sufficiently confident inputs stop before the final layer. Harder inputs continue through more layers. The result is input-dependent computation: the model can reduce average work without forcing every request to use a smaller network.

Artificial Intelligence 04 Sep 2026 10 min read

Gradient Noise in Mini-Batch Training

Neural network training usually updates model parameters from a small batch of examples rather than computing a gradient over the entire training set. That makes each update cheaper, but it also means the update direction depends on which examples happened to enter the batch. This variation is often called gradient noise. It is not necessarily a bug. It is a consequence of estimating a dataset-wide gradient from a sample, and it creates an important trade-off between computation per update, update variability, and training throughput.