Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 07 Sep 2026 9 min read

Weight Decay in Adam and AdamW

A training configuration can contain a parameter named weight_decay without making it obvious what operation the optimizer actually performs. That ambiguity matters most with adaptive optimizers such as Adam: adding an L2 penalty to the loss and directly decaying parameters are not generally the same update. The distinction is easy to miss because the two procedures are closely related under ordinary stochastic gradient descent (SGD). Once an optimizer rescales different coordinates using gradient history, however, the equivalence breaks.

Artificial Intelligence 07 Sep 2026 9 min read

Transfer Hyperparameters Across Model Width with MuP

Scaling a neural network creates an expensive tuning problem. A learning rate that works for a small prototype may behave differently after hidden dimensions become much wider. If every model size needs a fresh hyperparameter sweep, experimenting on small models saves less compute than it first appears. Maximal Update Parametrization, usually written MuP or μP, addresses this problem by changing how parameter initialization and learning rates scale with model width. The goal is not to make a wider model identical to a narrow one. It is to make important training dynamics behave consistently enough that hyperparameters tuned on a smaller proxy can often transfer to a wider target.

Artificial Intelligence 07 Sep 2026 10 min read

Train Neural Networks with Sharpness-Aware Minimization

A neural network can reach low training loss at parameter values where a small change to the weights makes the loss rise sharply. Standard optimization does not directly discourage this behavior: it mainly asks whether the loss is low at the current parameters. Sharpness-Aware Minimization (SAM) changes the training objective. Instead of optimizing only the loss at the current weights, it approximately optimizes the worst loss in a small neighborhood around them. The practical idea is simple: find a nearby parameter perturbation that makes the current mini-batch harder, then update the original model using the gradient measured at those perturbed parameters.

Artificial Intelligence 07 Sep 2026 10 min read

Resolve Conflicting Gradients in Multi-Task Learning

Training one neural network to solve several tasks can reduce duplicated computation and let related tasks share useful representations. It also creates a problem that single-task training does not have: two losses can ask the same shared parameter to move in opposing directions during the same update. Simply adding the losses does not make that disagreement disappear. Their gradients are added too, so one task can partially cancel another or dominate the shared update. Gradient surgery is a family of techniques that changes task gradients before combining them. A well-known example is projected conflicting gradients, commonly called PCGrad, which removes a conflicting component of one task’s gradient relative to another.

Artificial Intelligence 07 Sep 2026 10 min read

Repetition Controls for Language Model Generation

A language model can produce fluent text and still get stuck repeating a phrase, sentence pattern, or idea. A support reply may restate the same apology three times. A summarizer may loop over one point. A long generation can begin copying a short phrase again and again. The tempting fix is to turn up a “repetition” setting until the duplicate text disappears. That can solve one symptom while creating another: names become awkward, code becomes invalid, required terms disappear, or a model avoids repeating words that the task genuinely needs.

Artificial Intelligence 07 Sep 2026 10 min read

Probe Neural Network Representations with Linear Classifiers

A neural network can produce the right output while leaving an important engineering question unanswered: what information exists inside its intermediate representations? Suppose an image classifier predicts product categories. You may want to know whether an early layer already separates shapes, whether a later layer distinguishes categories, or whether a supposedly irrelevant attribute such as camera source remains easy to recover. Looking only at the final prediction does not answer those questions.

Artificial Intelligence 07 Sep 2026 9 min read

Prevent Catastrophic Forgetting in Continual Learning

Updating a neural network with new data sounds straightforward: continue training on the new examples and deploy the improved model. The difficulty is that an update which helps the new data can damage behavior the model learned earlier. A classifier that learns a new group of products, for example, may become worse at recognizing older groups even though those old classes never changed. This failure is called catastrophic forgetting. It matters most in continual learning, where a model learns from a sequence of tasks or data distributions instead of training once on a fixed mixed dataset.

Artificial Intelligence 07 Sep 2026 10 min read

Neural Network Pruning: Sparsity, Structure, and Speed

A neural network can contain parameters that contribute little to its useful predictions. Removing some of them can reduce storage or computation, but there is an important trap: a model with fewer nonzero weights is not automatically a model that runs faster. That distinction matters when developers use pruning to compress a trained network. The pruning rule determines what disappears, the hardware and runtime determine whether the resulting structure can be exploited, and the evaluation procedure determines whether the saved resources are worth any quality loss.

Artificial Intelligence 07 Sep 2026 12 min read

Measure Tokenization Efficiency Across Languages

Two prompts can communicate roughly the same amount of information and still consume very different numbers of model tokens. The difference can appear between languages, writing systems, domains, or even formatting styles. That matters because language-model systems usually operate on tokens rather than characters or words. A context window is measured in tokens. Many hosted APIs account for usage in tokens. Longer token sequences can also increase inference work, although the exact latency and compute effect depends on the model, serving stack, batching, caching, and whether the tokens belong to the input or generated output.

Artificial Intelligence 07 Sep 2026 9 min read

Knowledge Distillation for Smaller Classifiers

A model can be accurate enough for a product and still be too expensive to deploy. A large classifier may exceed a mobile memory budget, miss a latency target, or cost too much when every request requires substantial compute. Replacing it with a smaller model reduces those costs, but training the smaller model only from ground-truth labels can leave useful information behind. Knowledge distillation addresses this problem by training a smaller student model to learn from a stronger teacher model. Instead of seeing only the correct class, the student can also learn how the teacher distributes its confidence across the alternatives.

Artificial Intelligence 07 Sep 2026 9 min read

Improve Reasoning Reliability with Self-Consistency

A language model can reach different answers to the same reasoning problem depending on how generation unfolds. One sampled path may make an arithmetic mistake, another may misread a condition, and a third may reach the correct result. If an application trusts only one path, its answer depends heavily on that single generation. Self-consistency uses this variability instead of trying to eliminate it. It samples several reasoning paths for the same problem, extracts their final answers, and chooses the answer supported by the largest share of the samples. The technique is an inference-time strategy: it does not require changing model weights.

Artificial Intelligence 07 Sep 2026 13 min read

Handle Label Shift in Deployed Classifiers

A classifier can keep seeing familiar inputs and still make worse decisions after deployment. One reason is that the frequency of the classes has changed. Imagine a model trained to classify support tickets as billing, account, or technical. During training, billing tickets made up 20% of examples. After a pricing migration, billing issues temporarily rise to 50%. The model has not changed, but one part of the environment has: the prior probability of each class.

Artificial Intelligence 07 Sep 2026 12 min read

Gradient Noise Scale and Batch Size

Increasing a neural network’s batch size can make more accelerators useful, but the benefit does not grow indefinitely. At some point, processing more examples before each update gives a cleaner estimate of nearly the same gradient direction while consuming additional examples and compute. Gradient noise scale provides a useful mental model for this transition. It compares the variation in per-example gradients with the strength of their average. When gradient estimates are noisy relative to their mean, averaging more examples can remove meaningful noise. When they are already stable, a larger batch has less statistical work left to do.

Artificial Intelligence 07 Sep 2026 9 min read

Gradient Centralization in Neural Network Training

Neural network optimizers normally consume the gradients produced by backpropagation directly. Gradient centralization inserts one small transformation between those two steps: for selected weight tensors, it subtracts the mean of each gradient vector before the optimizer uses it. That operation is easy to implement, but its effect is easy to misunderstand. It is not gradient clipping, because it does not cap large values. It is not normalization, because it does not divide by a norm or standard deviation. It changes the direction of the update by removing one particular component.

Artificial Intelligence 07 Sep 2026 11 min read

Diversify RAG Retrieval with Maximal Marginal Relevance

A retrieval-augmented generation (RAG) system can retrieve highly relevant passages and still build a poor context. The problem appears when the top results repeat the same fact in slightly different wording. Suppose five retrieved chunks all explain how to reset an API token, while a lower-ranked chunk explains the permission change that must happen afterward. Filling the context window with the five near-duplicates gives the language model less useful evidence than selecting a smaller set that covers both parts of the task.

Artificial Intelligence 07 Sep 2026 11 min read

Diagnose Hubness in Embedding Retrieval

Embedding retrieval is usually explained one query at a time: encode the query, compare it with stored vectors, and return the nearest items. That view can hide a collection-level failure mode. A document may look reasonably similar to many unrelated queries and therefore appear in far more result lists than it should. Such an item is called a hub. The broader phenomenon, hubness, is a tendency for some points in a vector space to become nearest neighbors of unusually many other points. It matters because a retriever can have healthy-looking similarity scores while repeatedly wasting top positions on generic or geometrically favored items.

Artificial Intelligence 07 Sep 2026 10 min read

Diagnose and Prevent Dead ReLU Neurons

ReLU is a simple and effective activation function, but its simplicity creates a failure mode that is easy to miss. A neuron can reach a state where its pre-activation is negative for every relevant input. Its ReLU output is then always zero, and the gradient through that activation is also zero. If this persists, the neuron may stop participating in learning. This is commonly called a dead ReLU or dying ReLU problem. It does not mean that every zero activation is a defect: sparse activations are a normal consequence of ReLU. The useful question is whether a unit is inactive only for some inputs or effectively inactive across the data it needs to model.

Artificial Intelligence 07 Sep 2026 9 min read

Defer Uncertain Classifier Predictions with Selective Classification

A classifier does not have to make an automated decision for every input. In many systems, forcing a prediction on the hardest cases is exactly what creates expensive mistakes. Imagine a model that routes customer support messages. Clear password-reset requests can be handled automatically, while ambiguous messages could be sent to a human queue. The important design question is no longer only “How accurate is the classifier?” It is also “How accurate is it on the cases we allow it to handle?”

Artificial Intelligence 07 Sep 2026 11 min read

Compress Neural Network Layers with Low-Rank Factorization

Large neural networks spend much of their memory and computation multiplying activations by weight matrices. Some of those matrices contain more independent structure than the model actually needs for a particular deployment. If so, we can approximate one large matrix with two smaller matrices and reduce the number of stored parameters and multiply-add operations. This technique is called low-rank factorization. The central idea is simple, but using it well requires more than choosing a smaller number. Compression changes the weights, approximation error can accumulate through a network, and fewer arithmetic operations do not guarantee lower wall-clock latency on every device.

Artificial Intelligence 07 Sep 2026 11 min read

Chunking Documents for RAG Without Losing Context

A retrieval-augmented generation (RAG) system can have a strong embedding model and still retrieve poor evidence. One common reason is document chunking: the text was divided into units that are awkward to search or incomplete when read on their own. If chunks are too large, one embedding must represent several unrelated ideas and retrieval becomes less precise. If chunks are too small, the retrieved text may omit the definitions, qualifiers, or surrounding steps needed to answer correctly. The problem is therefore not to find one universally correct chunk size. It is to create retrieval units that are focused enough to match a query and complete enough to be useful after retrieval.

Artificial Intelligence 07 Sep 2026 10 min read

Average Neural Network Weights with Stochastic Weight Averaging

A neural network rarely finishes training at the only useful point in parameter space. Late in training, stochastic gradient descent (SGD) can visit several nearby parameter settings that all perform reasonably well, while the final checkpoint represents only one of them. Stochastic weight averaging (SWA) turns that observation into a simple training technique: collect model parameters from multiple points late in an SGD trajectory and compute their arithmetic mean. The result is one model with averaged weights, so inference does not require running an ensemble of all collected checkpoints.

Artificial Intelligence 07 Sep 2026 10 min read

Adapt Neural Networks with Gradient Reversal

A classifier can perform well in evaluation and then weaken after deployment because the inputs changed. Product photos may come from a new camera, support messages may use different vocabulary, or sensor readings may come from different hardware. Collecting labels for the new environment can be expensive even when unlabeled examples are easy to obtain. Domain-adversarial training addresses one version of this problem. It asks a feature extractor to support the prediction task while making the source and target domains difficult to distinguish. A gradient reversal layer makes those two goals trainable with ordinary backpropagation by reversing the domain classifier’s gradient before it reaches the feature extractor.

Artificial Intelligence 06 Sep 2026 10 min read

Use Best-of-N Sampling to Spend More Compute at Inference

A language model can produce several plausible answers to the same prompt. That variability is often treated as noise, but it can also be used deliberately: generate multiple candidates, evaluate them, and return the strongest one. This pattern is called best-of-N sampling. Instead of trusting one generation, the system samples N responses and uses a scoring rule to select one. The extra samples spend more compute at inference time in exchange for more opportunities to find a good response.

Artificial Intelligence 06 Sep 2026 9 min read

Token and Sequence Biases for LLM Decoding

Sometimes an LLM produces generally good text but makes one narrow decoding choice too often. Perhaps a domain-specific abbreviation should be preferred, a deprecated product name should be discouraged, or a particular token must not appear in generated text. Changing temperature is a poor fit for this problem because temperature affects the whole next-token distribution. Retraining a model is usually excessive when the desired change is local. Token and sequence biases provide a narrower tool: modify selected prediction scores during decoding while leaving the model parameters unchanged.