Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 10 Sep 2026 9 min read

Reduce Repetitive Generation with Unlikelihood Training

Reduce Repetitive Generation with Unlikelihood Training A language model can learn to predict ordinary text well and still assign too much probability to behavior you don’t want at generation time. Repetition is a common example: once a phrase appears, the model may keep making recently used tokens plausible enough that a decoding loop becomes hard to escape. Changing the decoder can hide some of this behavior, but it doesn’t change the probabilities learned by the model. Unlikelihood training takes a different approach. During training, it identifies undesirable candidates and explicitly pushes their probabilities down while the usual likelihood objective pushes desired tokens up.

Artificial Intelligence 10 Sep 2026 10 min read

Prevent FP16 Gradient Underflow with Dynamic Loss Scaling

Prevent FP16 Gradient Underflow with Dynamic Loss Scaling Mixed-precision training can reduce memory use and accelerate supported operations, but float16 introduces a numerical problem that is easy to miss: some gradients are too small to survive in FP16. They can round to zero before the optimizer gets a chance to use them. Loss scaling addresses that problem by multiplying the loss before backpropagation, which multiplies the resulting gradients by the same factor. The gradients are divided by that factor before the optimizer update, so the intended update is unchanged when the arithmetic remains finite. Dynamic loss scaling adjusts the factor during training so you don’t have to guess one fixed value for the whole run.

Artificial Intelligence 10 Sep 2026 9 min read

Prepare Models for Low-Precision Inference with Quantization-Aware Training

A model can work well in floating point and lose useful accuracy after its weights or activations are quantized for deployment. The problem is not mysterious: rounding and clipping change the numbers that flow through the network, while the original model was optimized without those changes in the loop. Quantization-aware training (QAT) exposes the model to an approximation of those low-precision numerics while its parameters can still adapt. Training remains differentiable in floating point, but the forward computation simulates the quantization errors expected after conversion.

Artificial Intelligence 10 Sep 2026 10 min read

Pack Training Sequences Without Leaking Between Examples

Pack Training Sequences Without Leaking Between Examples Language-model training often wastes computation on padding. If a batch contains examples with very different lengths, shorter examples are extended with padding so tensors have compatible shapes. The model still has to move those tensor positions through parts of the training pipeline even though they contain no training content. Sequence packing reduces that waste by placing multiple shorter examples into one fixed-length training sequence. The idea is simple; the boundary handling is not. If attention or loss masks are wrong, one example can accidentally use another example as context, or the model can be trained to predict tokens that should not count as targets.

Artificial Intelligence 10 Sep 2026 10 min read

Neural Collapse in Deep Classifiers

Neural Collapse in Deep Classifiers A classifier can keep changing after it already predicts every training example correctly. Cross-entropy loss can continue to fall, feature vectors can reorganize, and the final classification layer can become increasingly regular. Looking only at training accuracy hides all of that movement. Neural collapse is a name for a collection of geometric patterns that can emerge late in the training of deep classifiers. The striking part isn’t simply that examples from the same class become similar. Under the conditions where neural collapse appears, within-class variation can shrink while class centers and classifier weights approach a highly symmetric arrangement.

Artificial Intelligence 10 Sep 2026 8 min read

Mask Prompt Tokens During Instruction Fine-Tuning

A supervised language-model example often contains more than the text you want the model to produce. It may include a system message, a user request, separators, and an assistant answer. If you compute next-token loss over the entire sequence, the model is trained to predict all of those tokens, not just the assistant response. That may be intentional for some training objectives. For instruction fine-tuning, though, developers often want the prompt to provide context while only selected response tokens contribute to the supervised loss. A loss mask makes that distinction explicit.

Artificial Intelligence 10 Sep 2026 11 min read

Improve LLM Fine-Tuning with Rejection Sampling

Improve LLM Fine-Tuning with Rejection Sampling Suppose you can tell a good model response from a bad one, but writing thousands of ideal responses by hand is expensive. A capable language model may already produce acceptable answers some of the time. The problem is that those answers are mixed with weaker ones. Rejection sampling fine-tuning turns that observation into a data-generation loop. For each prompt, generate several candidate responses, evaluate them, keep responses that satisfy a selection rule, and use the accepted prompt-response pairs for supervised fine-tuning. The method can concentrate training on behavior you want without requiring a human to author every target from scratch.

Artificial Intelligence 10 Sep 2026 10 min read

Evaluate Knowledge Edits Before Trusting Model Updates

Changing one fact in a language model sounds simpler than retraining it. If a product name changes, an organization moves offices, or a fictional knowledge base is updated, knowledge editing aims to change a model’s behavior for that information without running broad training again. The hard part isn’t making one prompt produce the new answer. The hard part is knowing what else changed. A useful evaluation therefore asks more than “did the edit work?” It checks whether the new fact survives reasonable paraphrases, whether unrelated behavior stays stable, and whether the model can use the edited information when another answer depends on it. This article builds that evaluation model and shows how to turn it into a practical test suite.

Artificial Intelligence 10 Sep 2026 10 min read

Distill Sequence Models with Teacher-Generated Outputs

Distill Sequence Models with Teacher-Generated Outputs A large text generator may produce useful outputs but still be too expensive for the latency, memory, or throughput budget of a deployment. Training a smaller model on the original dataset is the obvious baseline, but it throws away information encoded in the larger model’s behavior. Sequence-level knowledge distillation offers another option: let a capable teacher generate target sequences, then train a smaller student to reproduce those sequences. The student learns from concrete examples of what the teacher tends to produce rather than only from the original human targets or from the teacher’s next-token probabilities.

Artificial Intelligence 09 Sep 2026 9 min read

Use Predictive Entropy to Detect Uncertain Classifications

Use Predictive Entropy to Detect Uncertain Classifications A classifier can return the same predicted label for two inputs while being much less certain about one of them. If an application only keeps the winning label, that difference disappears. Predictive entropy gives you a compact way to preserve it. It summarizes how spread out a classifier’s predicted probability distribution is: concentrated probability produces low entropy, while probability spread across several classes produces higher entropy.

Artificial Intelligence 09 Sep 2026 11 min read

Train with Unlabeled Data Using Mean Teacher

Labeled examples are often the expensive part of an AI system. You may have millions of inputs but only a small subset with trustworthy labels. Training only on the labeled subset ignores information in the rest of the data, while assigning guessed labels too aggressively can teach the model its own mistakes. Mean Teacher is a semi-supervised learning method for this situation. It trains a normal model, called the student, while maintaining a second model, called the teacher, whose parameters are an exponential moving average of the student’s parameters. The student learns from real labels when they exist and is also encouraged to make predictions that agree with the teacher on unlabeled inputs.

Artificial Intelligence 09 Sep 2026 12 min read

Train Discrete Neural Operations with Straight-Through Estimators

Neural networks are usually trained with gradient descent, which depends on small changes in parameters producing informative changes in the loss. A discrete operation can break that assumption. Rounding a value, choosing a binary gate, or selecting a quantized level may be exactly what the forward computation needs, yet its derivative can be zero almost everywhere or undefined at transition points. A straight-through estimator (STE) is a practical way to keep training in that situation. The forward pass uses the discrete operation, while the backward pass substitutes a simpler derivative so that a gradient can flow through it. The important consequence is easy to miss: the backward signal is generally not the true derivative of the discrete forward computation. It is a deliberately chosen surrogate.

Artificial Intelligence 09 Sep 2026 10 min read

Select Active Learning Examples with BALD

When labels are expensive, training on every available example may be impractical. An active learning system tries to spend its labeling budget selectively: train a model on the labels already available, score unlabeled examples, request labels for useful examples, then retrain. A common first idea is to label the examples with the highest predictive entropy. That can help, but entropy mixes together two different reasons for uncertainty. The model may be uncertain because it does not yet know enough, or because the input itself is genuinely ambiguous. More labels are most valuable for the first case.

Artificial Intelligence 09 Sep 2026 9 min read

LLM Sampling: Temperature, Top-K, and Top-P

A language model does not normally produce a single inevitable next token. Given a prefix, it assigns scores to many possible tokens. A decoding algorithm then decides how to turn those scores into the next output. That last step matters. If you sample too freely, a model can drift into unlikely continuations. If you restrict sampling too aggressively, outputs can become repetitive or lose useful variation. Parameters such as temperature, top-k, and top-p control different parts of this trade-off, so treating them as interchangeable “creativity settings” leads to confusing results.

Artificial Intelligence 09 Sep 2026 12 min read

Evaluate Sequence Models Beyond Training Length

A sequence model can score well on a random test split and still fail when an input is longer than the sequences it saw during training. This matters for language, symbolic reasoning, event sequences, and other tasks where production inputs do not have one fixed length. The problem is easy to hide. If training and test examples come from the same length distribution, an aggregate metric mostly measures performance on familiar lengths. It does not tell you whether the model learned a rule that extends to longer sequences or a strategy that works only inside the observed range.

Artificial Intelligence 09 Sep 2026 11 min read

Diversify RAG Retrieval with Maximum Marginal Relevance

A retrieval-augmented generation (RAG) system can retrieve highly relevant chunks and still build a poor context. The problem is redundancy. Imagine a support assistant answering a question about an API timeout. Vector search returns five chunks, but four are slightly different copies of the same timeout definition. The fifth useful chunk about retry behavior never reaches the model. Each result looked relevant in isolation, yet the set wastes most of its context budget repeating one idea.

Artificial Intelligence 09 Sep 2026 10 min read

Diagnose Embedding Anisotropy Before Tuning Vector Search

Diagnose Embedding Anisotropy Before Tuning Vector Search A vector search system can behave strangely even when its indexing code and similarity calculation are correct. Unrelated items may receive surprisingly high cosine similarities, score differences may look compressed, or many embeddings may point in broadly similar directions. One possible cause is embedding anisotropy: the vectors occupy some directions much more strongly than others instead of being distributed evenly through the representation space. Anisotropy is a property of the embedding geometry, not proof that retrieval is broken. The useful question is whether that geometry is hurting the decisions your system makes.

Artificial Intelligence 09 Sep 2026 9 min read

Diagnose and Prevent Dying ReLU Units

ReLU is one of the simplest neural-network activation functions: negative inputs become zero and positive inputs pass through unchanged. That simplicity makes optimization efficient, but it creates a failure mode that can quietly waste model capacity. A unit can move into a state where its pre-activation is negative for every relevant training example, so its ReLU output stays zero and the unit stops receiving a useful gradient through that activation.

Artificial Intelligence 09 Sep 2026 10 min read

Detect Representation Collapse in Self-Supervised Learning

Self-supervised learning can train an encoder without manually assigning a class label to every example. But removing labels also removes an obvious force that tells different examples to occupy meaningfully different parts of representation space. A badly designed objective can therefore admit a trivial solution: the encoder maps many or all inputs to essentially the same representation. This failure is called representation collapse. The training loss may even look good, because a model that emits the same vector for two augmented views of every input has achieved perfect agreement without learning useful distinctions.

Artificial Intelligence 09 Sep 2026 11 min read

Control Beam Search Length Bias with Length Normalization

Beam search is a common way to decode sequence models when choosing the most likely token at every step is too shortsighted. It keeps several partial candidates alive, expands them, and repeatedly retains the strongest alternatives. There is a subtle problem: the score used for a sequence usually accumulates one log-probability per generated token. Because token probabilities are at most 1, their log-probabilities are normally non-positive. Extending a sequence therefore tends to make its raw cumulative score smaller. When finished candidates of different lengths compete directly, this can create a preference for outputs that end too early.

Artificial Intelligence 09 Sep 2026 8 min read

Contextual Calibration for Few-Shot Classifiers

A language model can behave like a classifier without any parameter updates: give it a few labeled examples, present a new input, and ask it to choose a label. The surprising problem is that the answer can depend not only on the new input, but also on details such as the prompt wording, demonstration order, and label tokens. That creates a practical debugging trap. A prompt may appear to teach the task while also giving the model a baseline preference for one answer before meaningful input is considered.

Artificial Intelligence 09 Sep 2026 10 min read

Choose Consensus Outputs with Minimum Bayes Risk Decoding

A generative model can assign high probability to an output that is not the most useful answer for your application. This is especially visible when several different outputs are plausible: a translation can have multiple valid phrasings, a summary can emphasize different details, and a structured generator can produce several semantically similar candidates. Greedy decoding chooses locally likely tokens. Beam search searches for a high-probability sequence. Sampling gives you diverse candidates. None of those methods, by itself, asks a different question that is often closer to the application goal: which candidate agrees best with the distribution of plausible outputs?

Artificial Intelligence 08 Sep 2026 11 min read

Vanishing and Exploding Gradients in Deep Networks

A deep neural network can have enough capacity to solve a task and still fail to learn because useful training signals do not reach all of its layers. Parameters near the output may update normally while earlier layers receive gradients that are almost zero. In the opposite case, gradients can grow so large that one optimizer step destabilizes the model. These are the vanishing-gradient and exploding-gradient problems. They are not simply labels for “training is bad.” They describe what happens to derivatives as backpropagation repeatedly applies the chain rule through many transformations.

Artificial Intelligence 08 Sep 2026 9 min read

Use Test-Time Augmentation Without Hiding Model Errors

A model usually makes one prediction from one representation of an input. That is convenient, but the representation may contain accidental details that should not determine the answer. A product photo can be shifted a few pixels. A scanned digit can be slightly rotated. A crop can place the object closer to one edge than another. Test-time augmentation (TTA) asks the trained model to predict several valid transformations of the same input and then combines those predictions. The technique can make inference less dependent on one particular view, but it also increases compute and can make predictions worse when the transformations change information that matters to the label.