Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 14 Sep 2026 5 min read

Contrastive Decoding with an Amateur Model

A language model can assign high probability to a token for several reasons. Some reflect context-specific structure; others reflect generic tendencies that also appear in a weaker model. Contrastive decoding separates those signals by scoring candidate tokens with two models rather than one. The larger model acts as an expert. A smaller or otherwise weaker model acts as an amateur. Generation favors tokens that the expert supports more strongly relative to the amateur, subject to a plausibility constraint from the expert distribution.

Artificial Intelligence 14 Sep 2026 6 min read

Compare RMSNorm and Layer Normalization

Normalization layers can look interchangeable when their outputs have similar shapes, but their invariances are not the same. RMSNorm rescales an activation vector using its root mean square without first subtracting the vector mean. Layer normalization centers the vector and then rescales it using its variance. That missing centering operation is the central distinction. It changes which transformations of an activation vector disappear under normalization and which remain visible to the rest of the network.

Artificial Intelligence 14 Sep 2026 6 min read

Bound Extreme Logits with Soft Capping

A transformer can produce logits whose magnitudes grow far beyond the range needed to express a strong preference. Large attention scores can make a softmax distribution extremely concentrated, while large output logits can make token probabilities nearly one-hot. A hard clamp can bound those values, but it introduces a flat region with an abrupt derivative change at the threshold. Logit soft capping uses a smooth saturating function instead. One common form is:

Artificial Intelligence 14 Sep 2026 6 min read

Accumulate Gradients Across Microbatches

A training batch can exceed accelerator memory even when model parameters and optimizer state fit comfortably. Activations from the forward pass often account for a large part of the remaining footprint, and their memory cost grows with the number of examples processed together. Gradient accumulation splits a larger logical batch into smaller microbatches. Each microbatch runs its own forward and backward pass, but the optimizer waits until several backward passes have contributed to the parameter gradients. This reduces the activation memory required for any single pass without requiring an optimizer update after every microbatch.

Artificial Intelligence 14 Sep 2026 5 min read

Account for Exposure Bias in Autoregressive Generation

An autoregressive model can receive a clean prefix at every training position and still face a different input distribution during generation. Training commonly scores the next reference token while conditioning on earlier reference tokens. At inference time, the prefix contains the model’s own outputs instead. This mismatch is called exposure bias. It matters because an early generation error does more than make one token incorrect. That token becomes part of the context for later predictions, placing the model in a prefix state that may have been rare or absent during training.

Artificial Intelligence 13 Sep 2026 6 min read

Trade Activation Memory for Recomputation with Gradient Checkpointing

Training a deep neural network requires more memory than its parameters alone suggest. Backpropagation needs intermediate values from the forward pass, and retaining those activations across many layers can consume a large share of accelerator memory. Gradient checkpointing changes that storage policy. Instead of retaining every intermediate activation until its gradient is computed, training keeps selected boundary tensors and reconstructs omitted intermediates by running parts of the forward computation again during the backward pass. The model function need not change, but the execution schedule does.

Artificial Intelligence 13 Sep 2026 7 min read

Steer LLM Behavior with Activation Vectors

A transformer can produce different continuations without changing its prompt, weights, or decoding settings if an internal activation is modified during the forward pass. Activation steering uses this property as an inference-time control mechanism. A vector representing a target attribute is added to, or subtracted from, a hidden representation at selected model locations. The operation is simple, but its effect depends on where the vector came from, where it is injected, and how strongly it is scaled. A direction that separates two sets of prompts in one layer is not automatically a portable semantic control across layers, model revisions, or prompt distributions.

Artificial Intelligence 13 Sep 2026 6 min read

Reuse Shared Prefix State in LLM Inference

Autoregressive LLM serving often repeats the same initial tokens across many requests. A fixed system prompt, tool schema, or document prefix can occupy thousands of tokens before request-specific text begins. Computing attention state for that identical prefix on every request repeats prefill work that has already produced the same cached keys and values under compatible execution conditions. Prefix caching stores reusable attention state for such shared token prefixes. It changes the amount of prefill computation required for a cache hit, but it does not make arbitrary similar prompts interchangeable. The reusable unit is tied to exact model input state, not semantic resemblance.

Artificial Intelligence 13 Sep 2026 6 min read

Regularize Classifier Targets with Label Smoothing

A classifier trained with one-hot targets is rewarded for pushing the target class probability toward one and every other class probability toward zero. Cross-entropy keeps applying pressure in that direction even after the predicted class is already correct. Label smoothing changes that pressure by replacing the exact one-hot target with a distribution that reserves some mass for other classes. That small change affects more than the target tensor. It changes the gradient on every output logit, limits the incentive for extreme class separation, and alters how predicted probabilities should be interpreted.

Artificial Intelligence 13 Sep 2026 7 min read

Quantize KV Caches with Separate Key and Value Error Budgets

Autoregressive transformer inference keeps past key and value tensors so each new token can attend to prior positions without recomputing the full prefix. That KV cache grows with sequence length, layer count, batch size, and the number of stored key-value heads. At long contexts, its memory footprint can become a direct limit on concurrent requests or usable context length. Quantizing the KV cache reduces bytes per stored element. The resulting approximation is not equivalent to quantizing a passive data structure, however. Cached keys participate in attention score computation, while cached values are mixed according to the resulting attention weights. Error in those two tensors therefore enters the attention operation at different points.

Artificial Intelligence 13 Sep 2026 8 min read

Preserve Attention Sinks in Sliding-Window LLM Inference

A sliding-window KV cache seems mechanically simple: keep the most recent tokens, evict older key-value entries, and continue decoding within a fixed memory budget. The complication is that some transformer models place substantial attention mass on a small set of early positions even when those positions carry little direct semantic relevance to the current token. Those positions are often called attention sinks. If a cache policy removes them while preserving only the newest tokens, the attention distribution seen during decoding can change abruptly. A bounded cache can therefore behave differently from full-context inference even when the evicted text appears unrelated to the current request.

Artificial Intelligence 13 Sep 2026 6 min read

Preserve Attention Sinks in Bounded KV Caches

A decoder that keeps only the newest key-value states can degrade even when its cache still contains enough recent text for the immediate task. In some transformer models, early token positions attract substantial attention across later decoding steps. Evicting those states changes the attention distribution, not just the amount of accessible history. Attention sink retention addresses that specific failure mode. A bounded cache preserves a small prefix of initial key-value states together with a moving window of recent states. Tokens between those regions can be discarded, keeping cache size bounded as generation continues.

Artificial Intelligence 13 Sep 2026 7 min read

Merge Retrieval Rankings with Reciprocal Rank Fusion

A lexical retriever and an embedding retriever can return useful results for the same query while assigning scores that have no common numerical meaning. Adding those raw scores treats incomparable scales as if they were calibrated measurements. Reciprocal rank fusion avoids that assumption by combining positions rather than score magnitudes. This makes RRF useful in retrieval-augmented generation systems that mix distinct retrieval signals. Each retriever keeps its own scoring model. The fusion layer only needs ordered result lists and stable document identities.

Artificial Intelligence 13 Sep 2026 6 min read

Inspect Intermediate Transformer Predictions with Logit Lens

A transformer produces its next-token distribution only after the final block, yet every block updates the residual state that eventually feeds that prediction. Logit lens examines those intermediate states by mapping them through the model’s output path into vocabulary logits. The result is a sequence of provisional token distributions across model depth. The method is attractive because it reuses components already present in the model. Its output also needs careful interpretation. An intermediate residual state was not necessarily optimized to behave like a final residual state, so a readable token ranking is a diagnostic projection rather than a direct transcript of internal computation.

Artificial Intelligence 13 Sep 2026 5 min read

Focus Classification Loss with Focal Modulation

Cross-entropy gives every classified example a loss determined by the probability assigned to its target class. When a training batch contains many examples the model already classifies with high confidence, their individual losses may be small yet their aggregate contribution can still occupy a substantial part of the objective. Focal loss changes that balance with a confidence-dependent multiplier. The mechanism is not a new classifier head or sampling strategy. It modifies the loss so that examples with high target-class probability are attenuated more strongly than examples with low target-class probability.

Artificial Intelligence 13 Sep 2026 7 min read

Diversify Retrieval Results with Maximum Marginal Relevance

A retriever can fill its top positions with passages that are individually relevant but nearly interchangeable. Several chunks from one document may repeat the same fact, leaving little room for other evidence in a fixed context budget. Maximum marginal relevance, commonly abbreviated MMR, addresses this at the selection stage by considering both query relevance and redundancy with items already chosen. MMR does not change the embedding model or recover candidates that retrieval missed. It reranks a candidate pool. That boundary matters: the method can improve variety among available candidates, but it cannot compensate for poor candidate recall.

Artificial Intelligence 13 Sep 2026 7 min read

Diagnose Embedding Anisotropy in Vector Retrieval

A vector retriever can return stable nearest neighbors even when its embedding space uses only a narrow set of directions. In that case, high cosine similarity may reflect shared global structure as well as query-specific semantic alignment. This geometric concentration is commonly described as embedding anisotropy. Anisotropy matters at the retrieval boundary because nearest-neighbor search operates on the geometry it receives. An index can reproduce cosine or inner-product rankings correctly while those rankings still have weak separation between relevant and irrelevant candidates. Treating every retrieval issue as an indexing problem can therefore hide a representation problem upstream.

Artificial Intelligence 13 Sep 2026 6 min read

Detect Out-of-Distribution Inputs with Classifier Energy Scores

A classifier can assign high softmax confidence to an input that does not resemble the data used to fit its parameters. Softmax normalizes scores across the available classes; it does not add a separate class for unfamiliar inputs. As a result, a large maximum probability is not evidence that an input belongs to the expected data distribution. Energy-based out-of-distribution detection uses the full logit vector to produce a scalar score before a deployment policy decides whether an input looks familiar enough to accept. The score is simple to compute for an existing classifier, but its interpretation depends on the model, temperature, data regime, and threshold calibration.

Artificial Intelligence 13 Sep 2026 6 min read

Control Token Repetition with Logit Penalties

A language model can assign high probability to a token that has already appeared several times in the generated text. If the decoder keeps selecting that token or a short pattern containing it, the output may settle into repetition even though each individual choice is plausible under the model. A repetition penalty changes this behavior at decoding time. It modifies candidate scores according to token history before the next token is selected. The model parameters stay fixed, but the effective distribution used by the decoder no longer matches the model’s unmodified next-token distribution.

Artificial Intelligence 13 Sep 2026 7 min read

Control Sequence Length Bias in Beam Search Scoring

Beam search keeps several partial sequences alive while decoding, but the score used to compare those sequences can create a systematic preference for particular lengths. With the common sum of token log probabilities, each additional token contributes a value at or below zero. A completed sequence can therefore lose score simply by continuing, even when the continuation is plausible. This is not only a property of beam width. It comes from the objective used to rank hypotheses. Changing the beam size changes how much of the search space is explored; changing the scoring rule changes which sequences the search considers preferable.

Artificial Intelligence 13 Sep 2026 7 min read

Control Sequence Length Bias in Beam Search

Beam search can return a shorter sequence even when a longer candidate contains locally plausible tokens at every position. The behavior follows directly from sequence scoring: autoregressive models multiply conditional token probabilities, or equivalently add their log probabilities. Since token probabilities are at most one, each additional token contributes a non-positive log term. That arithmetic makes sequence length part of decoding. Beam width changes which candidates survive, but it does not remove the scoring effect. A decoder therefore needs a deliberate policy for comparing hypotheses of different lengths and for deciding when a completed hypothesis is good enough to stop the search.

Artificial Intelligence 13 Sep 2026 6 min read

Contrastive Decoding with Expert and Amateur Models

A language model can give high next-token probability to text that is fluent but generic. Contrastive decoding changes the ranking by asking for a second signal: does a weaker model also find the same candidate easy to predict? A candidate favored by the expert but not by the amateur receives stronger relative support than one both models score highly. This is an inference-time mechanism. It does not alter either model’s parameters, and it does not convert the amateur model into a verifier. The decoder combines two token distributions and then selects from the resulting scores.

Artificial Intelligence 13 Sep 2026 7 min read

Contrast Expert and Amateur Models During Decoding

A language model can assign high probability to a token for two different reasons: the token may fit the prompt particularly well, or it may simply be common under many contexts. Contrastive decoding tries to separate those effects by comparing the next-token distributions of two models. A stronger expert supplies the main distribution, while a weaker amateur supplies a signal for patterns that do not require the expert’s extra capability.

Artificial Intelligence 13 Sep 2026 6 min read

Compare RMSNorm and LayerNorm in Transformers

LayerNorm and RMSNorm can occupy the same structural position in a transformer while applying different operations to the residual stream. LayerNorm subtracts the feature mean before scaling by a measure of spread. RMSNorm skips the centering operation and scales directly from the root mean square of the features. That small algebraic difference changes which transformations of an activation vector are removed by normalization. It also means that replacing one operation with the other is not, in general, a function-preserving edit to an existing model.