Skip to content

Archive

Training

24 articles
Artificial Intelligence 17 Sep 2026 6 min read

Teacher Forcing Creates a Prefix Distribution Gap

Autoregressive models predict the next token from a prefix. During teacher-forced training, that prefix usually comes from the reference sequence. During generation, it contains tokens emitted by the model itself. A prediction error can therefore change the context used for every later prediction. This difference is often called exposure bias. The useful engineering detail is more specific: training and generation can present different prefix distributions to the same conditional predictor. Token-level validation on clean reference prefixes does not fully characterize behavior after the model enters a prefix that its training data rarely presented.

Artificial Intelligence 17 Sep 2026 6 min read

Bound Gradient Updates with Global Norm Clipping

A single optimization step can contain gradients whose combined magnitude is far larger than the surrounding steps. If those gradients are passed directly to an optimizer, the resulting parameter update can move the model into a very different region of parameter space. Global norm clipping places a bound on the gradient magnitude before the optimizer consumes it. The mechanism is simple, but its behavior is easy to misread. It does not cap every gradient element independently, and it does not guarantee a fixed parameter-update norm for adaptive optimizers. It rescales the collected gradient vector when a chosen norm crosses a threshold.

Artificial Intelligence 16 Sep 2026 6 min read

Mask Padding Tokens in Language Model Loss

Variable-length text batches are commonly padded into rectangular tensors. The extra positions simplify batching, but they are not ordinary training targets. If padded target positions contribute to cross-entropy, the optimizer receives gradients for synthetic symbols that were introduced only to align tensor shapes. Preventing that signal requires a loss mask. An attention mask can stop selected positions from participating in attention, but that does not by itself remove their target terms from the objective.

Artificial Intelligence 16 Sep 2026 6 min read

Account for Exposure Bias in Autoregressive Decoding

An autoregressive model can receive cleaner context during training than it receives during generation. Under teacher forcing, the next-token prediction is conditioned on a reference prefix from the training sequence. During free-running decoding, the model instead conditions on tokens it generated itself. Once a generated token differs from the intended continuation, later predictions operate on a prefix that training may have represented less often. This mismatch is commonly called exposure bias. It is not simply a claim that autoregressive models make errors. The specific issue is that the distribution of prefixes presented to the model can change between optimization and generation, and an early deviation can change every subsequent conditional prediction.

Artificial Intelligence 14 Sep 2026 6 min read

Track Model Weights with an Exponential Moving Average

Optimizer updates can move model parameters back and forth even when the broader trajectory changes more gradually. An exponential moving average, or EMA, keeps a second parameter state that follows those updates with smoothing. The trainable model still receives ordinary optimizer updates; the EMA state is a derived copy used separately, often for evaluation or export. The mechanism is compact, but its behavior depends on decay, update frequency, initialization, and which state is actually saved or evaluated.

Artificial Intelligence 14 Sep 2026 7 min read

Reduce Repetition with Unlikelihood Training

An autoregressive language model is usually trained to increase the probability of the observed next token. That positive objective does not directly state which plausible but unwanted alternatives should receive less probability. When repetitive tokens or phrases remain locally probable, ordinary next-token training can leave generation with a strong route back into content that has already appeared. Unlikelihood training adds a negative signal for selected candidates. Instead of only rewarding the target token, the objective can also penalize tokens chosen because they represent an unwanted behavior, such as repetition within the generated prefix.

Artificial Intelligence 14 Sep 2026 5 min read

Account for Exposure Bias in Autoregressive Generation

An autoregressive model can receive a clean prefix at every training position and still face a different input distribution during generation. Training commonly scores the next reference token while conditioning on earlier reference tokens. At inference time, the prefix contains the model’s own outputs instead. This mismatch is called exposure bias. It matters because an early generation error does more than make one token incorrect. That token becomes part of the context for later predictions, placing the model in a prefix state that may have been rare or absent during training.

Artificial Intelligence 13 Sep 2026 6 min read

Trade Activation Memory for Recomputation with Gradient Checkpointing

Training a deep neural network requires more memory than its parameters alone suggest. Backpropagation needs intermediate values from the forward pass, and retaining those activations across many layers can consume a large share of accelerator memory. Gradient checkpointing changes that storage policy. Instead of retaining every intermediate activation until its gradient is computed, training keeps selected boundary tensors and reconstructs omitted intermediates by running parts of the forward computation again during the backward pass. The model function need not change, but the execution schedule does.

Artificial Intelligence 13 Sep 2026 6 min read

Clip Gradient Norms Before Optimizer Updates

A training update can become dominated by a gradient whose magnitude is far larger than the range seen in nearby iterations. Global norm clipping places a bound on that update signal before the optimizer consumes it. The operation is simple, but its behavior depends on what is included in the norm, where clipping occurs, and how it interacts with gradient accumulation and mixed-precision scaling. Norm clipping does not repair the source of unstable gradients. It changes the vector passed to the optimizer when its norm exceeds a chosen threshold.

Artificial Intelligence 12 Sep 2026 6 min read

Pack Transformer Training Sequences Without Cross-Sample Attention

Transformer batches often waste token slots on padding when examples have uneven lengths. Sequence packing reduces that waste by placing several shorter samples into one fixed-length token block. The arithmetic is attractive, but concatenation alone changes the training problem: tokens from one sample can attend to tokens from another unless the packed representation preserves sample boundaries. A correct packing scheme therefore has two jobs. It must fill token capacity more densely, and it must keep the model’s effective computation consistent with the intended independence of the original samples.

Artificial Intelligence 11 Sep 2026 11 min read

Clip Gradients Relative to Parameter Scale with AGC

Clip Gradients Relative to Parameter Scale with AGC A gradient can be large in absolute terms without being large for the parameter it updates. A gradient norm of 0.1 is modest next to a parameter norm of 10, but enormous next to a parameter norm of 0.001. Ordinary gradient clipping doesn’t see that distinction: it compares gradients with a fixed threshold. Adaptive gradient clipping (AGC) uses a different reference point. It compares a gradient’s norm with the norm of the parameter unit that gradient will update. If the gradient is too large relative to the parameter, AGC rescales it before the optimizer step.

Artificial Intelligence 08 Sep 2026 11 min read

Vanishing and Exploding Gradients in Deep Networks

A deep neural network can have enough capacity to solve a task and still fail to learn because useful training signals do not reach all of its layers. Parameters near the output may update normally while earlier layers receive gradients that are almost zero. In the opposite case, gradients can grow so large that one optimizer step destabilizes the model. These are the vanishing-gradient and exploding-gradient problems. They are not simply labels for “training is bad.” They describe what happens to derivatives as backpropagation repeatedly applies the chain rule through many transformations.

Artificial Intelligence 08 Sep 2026 10 min read

Train Neural Networks with Curriculum Learning

Most training pipelines treat the dataset as a fixed pool and repeatedly shuffle it. That is a strong default: it is simple, exposes the model to the full data distribution, and avoids assumptions about which examples should come first. But some learning problems have a useful notion of progression. A model may learn basic patterns more reliably before it is asked to handle noisy, ambiguous, or structurally difficult examples. Curriculum learning makes that progression explicit. Instead of changing the model architecture or the loss, it changes which training examples are emphasized at different stages of training. A common curriculum begins with easier examples and gradually introduces harder ones.

Artificial Intelligence 08 Sep 2026 11 min read

Trade Compute for Memory with Activation Checkpointing

A neural network can fit comfortably in accelerator memory for inference and still run out of memory during training. The reason is that training needs more than model weights. Backpropagation also needs intermediate values from the forward pass, and those activations can consume a large share of memory in deep models or with long sequences and large batches. Activation checkpointing reduces that memory pressure by deliberately not keeping every intermediate activation. Instead, training saves selected checkpoints and recomputes missing forward-pass values when the backward pass needs them. The trade is straightforward: keep fewer activations in memory, but perform extra computation.

Artificial Intelligence 08 Sep 2026 10 min read

Dynamical Isometry in Neural Networks

A deep neural network can look well scaled one layer at a time and still be difficult to optimize. Signals pass through many transformations, and small expansions or contractions can multiply with depth. By the time a gradient travels through the whole network, some directions may have nearly disappeared while others have been amplified dramatically. Dynamical isometry gives a precise way to reason about this problem. Instead of asking only whether the average gradient magnitude is reasonable, it asks how the network transforms different directions in its input space. The relevant object is the input-output Jacobian, and its singular values provide the main measurements.

Artificial Intelligence 07 Sep 2026 10 min read

Diagnose and Prevent Dead ReLU Neurons

ReLU is a simple and effective activation function, but its simplicity creates a failure mode that is easy to miss. A neuron can reach a state where its pre-activation is negative for every relevant input. Its ReLU output is then always zero, and the gradient through that activation is also zero. If this persists, the neuron may stop participating in learning. This is commonly called a dead ReLU or dying ReLU problem. It does not mean that every zero activation is a defect: sparse activations are a normal consequence of ReLU. The useful question is whether a unit is inactive only for some inputs or effectively inactive across the data it needs to model.

Artificial Intelligence 07 Sep 2026 10 min read

Average Neural Network Weights with Stochastic Weight Averaging

A neural network rarely finishes training at the only useful point in parameter space. Late in training, stochastic gradient descent (SGD) can visit several nearby parameter settings that all perform reasonably well, while the final checkpoint represents only one of them. Stochastic weight averaging (SWA) turns that observation into a simple training technique: collect model parameters from multiple points late in an SGD trajectory and compute their arithmetic mean. The result is one model with averaged weights, so inference does not require running an ensemble of all collected checkpoints.

Artificial Intelligence 06 Sep 2026 11 min read

Contrastive Learning for Text Embeddings

A text embedding model turns text into a vector so that software can compare meaning numerically. The difficult part is not producing vectors. A neural network can produce vectors for almost any input. The difficult part is teaching the geometry of those vectors so that distances correspond to the relationships your application cares about. Contrastive learning provides a practical way to do that. Instead of asking a model to predict a class label, you show it examples that should be close together and examples that should be farther apart. Training adjusts the encoder so that those relationships become easier to recover from the resulting vectors.

Artificial Intelligence 05 Sep 2026 10 min read

Stop Neural Network Training with Early Stopping

Neural network training does not automatically become more useful because it runs for more epochs. Training loss can continue falling while performance on unseen data stops improving or begins to deteriorate. That creates a practical question: when should training stop? A fixed epoch count is easy to configure, but it cannot know whether a particular run converged early or still needs more optimization. Early stopping answers this by monitoring performance on validation data during training. When the chosen validation metric stops improving for long enough, training ends. Used carefully, it can reduce wasted computation and limit unnecessary overfitting while preserving the checkpoint that performed best on validation data.

Artificial Intelligence 05 Sep 2026 11 min read

Evaluate Neural Networks with Exponential Moving Average Weights

Neural network parameters do not move smoothly toward a final solution. Stochastic optimization updates them using noisy minibatch gradients, so the weights used after one training step can differ slightly from those used after the next. Saving only the final step therefore makes one particular point on that training path responsible for evaluation and deployment. An exponential moving average, or EMA, keeps a second copy of the parameters that changes more gradually. Instead of replacing this copy with every new set of training weights, each update blends the previous average with the current parameters.

Artificial Intelligence 04 Sep 2026 11 min read

Use Label Smoothing Without Hiding Classification Mistakes

A classifier trained with ordinary cross-entropy is usually given a hard target: the correct class has probability 1, and every other class has probability 0. For a three-class example: cat dog bird 1.0 0.0 0.0 That target is simple and often appropriate. But it also asks the model to keep increasing the correct-class logit relative to the others even after the prediction is already very confident.

Artificial Intelligence 04 Sep 2026 8 min read

Use Dropout Without Breaking Neural Network Inference

A neural network can fit its training data well while performing poorly on new examples. One way to reduce this kind of overfitting is dropout, a training technique that randomly removes some activations on each forward pass. The idea is simple, but one detail causes many implementation bugs: dropout is intentionally stochastic during training and normally disabled during inference. If those modes are confused, evaluation becomes noisy or predictions use the wrong activation scale.

Artificial Intelligence 04 Sep 2026 9 min read

Batch Normalization in Neural Networks

A neural network can become harder to train when the scale and distribution of intermediate activations change as earlier layers update. One technique for controlling those activations is batch normalization, usually shortened to BatchNorm. BatchNorm looks simple: normalize a layer’s activations, then learn a scale and offset. The important detail is that its behavior depends on mode. During training it normally uses statistics from the current mini-batch. During inference it normally uses running statistics collected during training. Confusing those two paths can produce a model that trains normally but behaves poorly when deployed.

Artificial Intelligence 03 Sep 2026 10 min read

Stabilize Neural Network Training with Gradient Clipping

Neural network training can look healthy for many steps and then suddenly become unstable. The loss may jump, parameters may receive an unusually large update, or numerical values may become non-finite. One possible cause is an exploding gradient: the gradient becomes large enough that the resulting optimization step is destructive. Gradient clipping puts a limit on gradients before the optimizer uses them. It is especially useful when occasional gradient spikes are expected, but it is not a general repair for a bad learning rate, broken data, or an incorrect training loop.