Skip to content

Archive

Model Training

55 articles
Artificial Intelligence 10 Sep 2026 9 min read

Reduce Repetitive Generation with Unlikelihood Training

Reduce Repetitive Generation with Unlikelihood Training A language model can learn to predict ordinary text well and still assign too much probability to behavior you don’t want at generation time. Repetition is a common example: once a phrase appears, the model may keep making recently used tokens plausible enough that a decoding loop becomes hard to escape. Changing the decoder can hide some of this behavior, but it doesn’t change the probabilities learned by the model. Unlikelihood training takes a different approach. During training, it identifies undesirable candidates and explicitly pushes their probabilities down while the usual likelihood objective pushes desired tokens up.

Artificial Intelligence 10 Sep 2026 10 min read

Prevent FP16 Gradient Underflow with Dynamic Loss Scaling

Prevent FP16 Gradient Underflow with Dynamic Loss Scaling Mixed-precision training can reduce memory use and accelerate supported operations, but float16 introduces a numerical problem that is easy to miss: some gradients are too small to survive in FP16. They can round to zero before the optimizer gets a chance to use them. Loss scaling addresses that problem by multiplying the loss before backpropagation, which multiplies the resulting gradients by the same factor. The gradients are divided by that factor before the optimizer update, so the intended update is unchanged when the arithmetic remains finite. Dynamic loss scaling adjusts the factor during training so you don’t have to guess one fixed value for the whole run.

Artificial Intelligence 10 Sep 2026 9 min read

Prepare Models for Low-Precision Inference with Quantization-Aware Training

A model can work well in floating point and lose useful accuracy after its weights or activations are quantized for deployment. The problem is not mysterious: rounding and clipping change the numbers that flow through the network, while the original model was optimized without those changes in the loop. Quantization-aware training (QAT) exposes the model to an approximation of those low-precision numerics while its parameters can still adapt. Training remains differentiable in floating point, but the forward computation simulates the quantization errors expected after conversion.

Artificial Intelligence 10 Sep 2026 10 min read

Pack Training Sequences Without Leaking Between Examples

Pack Training Sequences Without Leaking Between Examples Language-model training often wastes computation on padding. If a batch contains examples with very different lengths, shorter examples are extended with padding so tensors have compatible shapes. The model still has to move those tensor positions through parts of the training pipeline even though they contain no training content. Sequence packing reduces that waste by placing multiple shorter examples into one fixed-length training sequence. The idea is simple; the boundary handling is not. If attention or loss masks are wrong, one example can accidentally use another example as context, or the model can be trained to predict tokens that should not count as targets.

Artificial Intelligence 10 Sep 2026 10 min read

Neural Collapse in Deep Classifiers

Neural Collapse in Deep Classifiers A classifier can keep changing after it already predicts every training example correctly. Cross-entropy loss can continue to fall, feature vectors can reorganize, and the final classification layer can become increasingly regular. Looking only at training accuracy hides all of that movement. Neural collapse is a name for a collection of geometric patterns that can emerge late in the training of deep classifiers. The striking part isn’t simply that examples from the same class become similar. Under the conditions where neural collapse appears, within-class variation can shrink while class centers and classifier weights approach a highly symmetric arrangement.

Artificial Intelligence 10 Sep 2026 8 min read

Mask Prompt Tokens During Instruction Fine-Tuning

A supervised language-model example often contains more than the text you want the model to produce. It may include a system message, a user request, separators, and an assistant answer. If you compute next-token loss over the entire sequence, the model is trained to predict all of those tokens, not just the assistant response. That may be intentional for some training objectives. For instruction fine-tuning, though, developers often want the prompt to provide context while only selected response tokens contribute to the supervised loss. A loss mask makes that distinction explicit.

Artificial Intelligence 09 Sep 2026 11 min read

Train with Unlabeled Data Using Mean Teacher

Labeled examples are often the expensive part of an AI system. You may have millions of inputs but only a small subset with trustworthy labels. Training only on the labeled subset ignores information in the rest of the data, while assigning guessed labels too aggressively can teach the model its own mistakes. Mean Teacher is a semi-supervised learning method for this situation. It trains a normal model, called the student, while maintaining a second model, called the teacher, whose parameters are an exponential moving average of the student’s parameters. The student learns from real labels when they exist and is also encouraged to make predictions that agree with the teacher on unlabeled inputs.

Artificial Intelligence 08 Sep 2026 10 min read

Use Self-Conditioning in Diffusion Models

A diffusion model repeatedly turns a noisy state into a cleaner one. Each denoising call normally receives the current noisy sample, a noise level or timestep, and any external condition such as a text embedding. Yet the previous call has already produced useful information about what the clean sample may look like. Throwing that estimate away means the next call must reconstruct similar information again from the new noisy state.

Artificial Intelligence 08 Sep 2026 9 min read

Stabilize Neural Network Weights with Exponential Moving Averages

A neural network’s parameters rarely move smoothly toward their final values. Mini-batch training produces noisy updates: one batch may push a weight in one direction, while the next pushes it partly back. The final checkpoint therefore represents one point on a noisy training path, not necessarily the most useful point near the end of that path. An exponential moving average (EMA) of model weights keeps a second set of parameters that changes more slowly than the actively trained model. Recent training states contribute more than old ones, but no single update immediately replaces the averaged weights.

Artificial Intelligence 08 Sep 2026 9 min read

Regularize Residual Networks with Stochastic Depth

Deep residual networks can overfit even when their skip connections make optimization manageable. Standard dropout can regularize individual activations, but residual architectures offer another useful unit to randomize: the entire residual branch. Stochastic depth randomly removes selected residual branches during training while keeping the skip path intact. A training example may therefore pass through a slightly shallower effective network on one step and the full set of blocks on another. At inference time, every residual branch is normally active.

Artificial Intelligence 08 Sep 2026 11 min read

Filter Synthetic Training Data with Rejection Sampling

Generating synthetic examples is easy; generating synthetic examples that are worth training on is harder. A language model can produce thousands of candidate answers, but blindly adding them to a training set can reinforce factual errors, weak reasoning, unwanted style, or artifacts of the generator itself. Rejection sampling provides a simple mental model for controlling that pipeline: generate one or more candidates, evaluate each candidate with an acceptance rule, and keep only candidates that pass. The acceptance rule might use deterministic checks, a learned reward model, another language model, human review, or a combination of signals.

Artificial Intelligence 07 Sep 2026 9 min read

Transfer Hyperparameters Across Model Width with MuP

Scaling a neural network creates an expensive tuning problem. A learning rate that works for a small prototype may behave differently after hidden dimensions become much wider. If every model size needs a fresh hyperparameter sweep, experimenting on small models saves less compute than it first appears. Maximal Update Parametrization, usually written MuP or μP, addresses this problem by changing how parameter initialization and learning rates scale with model width. The goal is not to make a wider model identical to a narrow one. It is to make important training dynamics behave consistently enough that hyperparameters tuned on a smaller proxy can often transfer to a wider target.

Artificial Intelligence 06 Sep 2026 10 min read

Stabilize Neural Network Evaluation with Exponential Moving Average Weights

A neural network’s final training step is not necessarily its most useful checkpoint. Stochastic optimization keeps moving the parameters as it follows noisy mini-batch gradients, so two nearby checkpoints can behave slightly differently even when training is otherwise healthy. An exponential moving average (EMA) of model weights gives you a second set of parameters that changes more smoothly. Instead of evaluating only the latest training weights, you maintain a weighted history in which recent weights matter most and older weights gradually fade away.

Artificial Intelligence 06 Sep 2026 11 min read

Sequence Parallelism for Lower Transformer Activation Memory

Large transformer training can run out of accelerator memory even after the model’s weights are split across several devices. The reason is easy to miss: tensor parallelism can shard expensive matrix multiplications while some intermediate activations remain replicated on every worker in the tensor-parallel group. Sequence parallelism removes part of that replication. For operations that work independently on each token, it partitions activations along the sequence dimension so each tensor-parallel worker keeps only a slice of the tokens. The workers temporarily reconstruct or reduce data where the tensor-parallel computation requires communication, then return to sequence-sharded activations.

Artificial Intelligence 06 Sep 2026 10 min read

Direct Preference Optimization for LLM Alignment

Supervised fine-tuning can teach a language model to imitate good answers, but many alignment problems are easier to express as comparisons: given two responses to the same prompt, which one is better? A preference dataset captures that signal as triples containing a prompt, a preferred response, and a rejected response. The challenge is turning those comparisons into model updates without treating a subjective preference as an ordinary next-token target. Direct Preference Optimization (DPO) provides one practical answer. It trains a policy model to increase its relative preference for chosen responses over rejected responses while measuring that change against a fixed reference model. Unlike a common reinforcement-learning-from-human-feedback pipeline, standard DPO does not require training a separate reward model and then running a reinforcement-learning optimizer.

Artificial Intelligence 05 Sep 2026 10 min read

Use Mixup to Regularize Neural Network Classifiers

A neural network can fit its training examples very well while learning decision boundaries that behave poorly between them. Ordinary augmentation helps by creating plausible variations of individual examples, but there is another useful idea: train the model on points that lie between pairs of examples. Mixup does this by interpolating both the inputs and their targets. If one image is labeled cat and another is labeled dog, mixup can create a synthetic input that is partly each image and a target that is partly each class. The model is then trained to produce a correspondingly mixed prediction.

Artificial Intelligence 05 Sep 2026 9 min read

Use Class-Weighted Loss for Imbalanced Classification

A classifier trained on imbalanced data can achieve a low average loss while learning the minority class poorly. If 99% of training examples belong to one class, errors on the remaining 1% contribute relatively little to an unweighted objective simply because they occur less often. Class-weighted loss changes that training signal. Instead of treating every example’s loss equally, it gives examples from selected classes more influence on parameter updates. This is useful when class frequency and the importance of learning each class are badly misaligned.

Artificial Intelligence 05 Sep 2026 10 min read

Reduce Transformer Padding with Length Bucketing

Transformer training often starts with a simple batching rule: shuffle the examples, take the next B sequences, and pad every sequence in the batch to the length of the longest one. The rule is correct, but it can waste substantial computation when sequence lengths vary widely. A batch containing a 900-token document and several 100-token documents must usually represent every sequence with 900 token positions. Attention masks prevent padding from acting like real input, but they do not necessarily make the padded positions free to process.

Artificial Intelligence 05 Sep 2026 10 min read

Pack Training Sequences to Reduce Padding Waste

Language-model training often processes sequences in fixed-size tensors. When examples have very different lengths, padding makes those tensors easy to batch but can leave many token positions doing little useful work. A batch that physically contains 8,000 positions may contain far fewer than 8,000 real training tokens. Sequence packing reduces this waste by placing multiple shorter examples into the same fixed-length training sequence. The idea is simple; the semantics are not. If packing accidentally lets one example attend to another, predicts across boundaries that should be independent, or assigns incorrect position IDs, the training objective changes rather than merely becoming more efficient.

Artificial Intelligence 04 Sep 2026 8 min read

Weight Decay in Neural Network Training

A neural network can keep reducing its training loss while learning parameter values that generalize poorly. Weight decay is one way to regularize training: it applies a small pressure that shrinks selected parameters as optimization proceeds. The idea sounds similar to adding an L2 penalty to the loss, and for plain stochastic gradient descent the two can be made equivalent by matching their scaling. With adaptive optimizers such as Adam, however, adding an L2 penalty to the gradient and directly decaying the weights are not generally the same operation. That distinction is why optimizers such as AdamW use decoupled weight decay.

Artificial Intelligence 04 Sep 2026 9 min read

Use Padding Masks for Variable-Length Transformer Batches

Transformer inputs rarely have identical lengths. One sentence may contain 8 tokens while another contains 30, yet efficient training and inference usually process multiple sequences in rectangular tensors. The usual solution is to add padding tokens to shorter sequences until their shapes match. Padding solves the shape problem but creates a semantic one: the added positions are not real input. If the model treats them like ordinary tokens, they can influence attention, pooling, and training loss. A padding mask tells the computation which positions are valid and which exist only to make the batch rectangular.

Artificial Intelligence 04 Sep 2026 9 min read

Stabilize Neural Network Evaluation with EMA Weights

Neural network training does not usually move parameters smoothly toward one final point. Mini-batch gradients are noisy, learning-rate schedules change step sizes, and later updates can move a model between nearby parameter settings with noticeably different validation results. That creates a practical question: should deployment use the parameters from one particular training step, or a smoothed version of several recent parameter states? An exponential moving average, or EMA, provides the second option. During training, it maintains a separate copy of the model parameters that changes more slowly than the actively optimized parameters. The optimizer still trains the ordinary model. The EMA copy is typically used for evaluation or inference.

Artificial Intelligence 04 Sep 2026 8 min read

Reduce Training Memory with Gradient Checkpointing

Training a neural network can run out of accelerator memory even when the model parameters fit comfortably. The missing piece is often activations: intermediate values produced during the forward pass and retained because backpropagation needs them later. Gradient checkpointing, also called activation checkpointing, trades extra computation for lower activation memory. Instead of keeping every intermediate activation until the backward pass, training keeps selected checkpoints and recomputes missing forward values when their gradients are needed.

Artificial Intelligence 04 Sep 2026 9 min read

Handle Class Imbalance in Machine Learning

A classifier can achieve impressive accuracy while failing on the cases you care about most. If only 1% of transactions are fraudulent, a model that predicts “not fraud” for every transaction is 99% accurate and still useless for detecting fraud. This is the practical problem of class imbalance: some target classes appear much less often than others. Imbalance does not automatically make a dataset bad, and it does not imply that every model needs special treatment. It does mean that accuracy can hide important errors and that the training objective may give rare examples too little influence.