Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 15 Sep 2026 6 min read

Treat Embedding Model Changes as Index Migrations

An embedding model change does not merely replace a function that emits arrays of the same length. It can change the coordinate system in which stored items and incoming queries are represented. Even when two models produce vectors with identical dimensions, their coordinates and similarity-score distributions are not interchangeable by default. That makes an embedding model version part of the index schema. A retrieval system that changes the query encoder while retaining vectors produced by an older encoder can still return numeric scores, but those scores no longer have a justified geometric interpretation unless cross-version compatibility is an explicit property of the models.

Artificial Intelligence 15 Sep 2026 6 min read

Steer Transformer Activations with Residual Stream Vectors

A transformer can produce different output behavior even when its weights and input tokens stay fixed. One way to cause that change is to alter an intermediate hidden state during the forward pass. Activation steering does this deliberately by adding a vector to a selected residual-stream position or set of positions. The mechanism is simple enough to express as an intervention, but its effect is not a global model setting. The chosen direction, coefficient, layer, token positions, and decoding setup all affect the result. Treating those choices as part of the inference configuration makes the behavior easier to reason about and test.

Artificial Intelligence 15 Sep 2026 6 min read

Quantize KV Caches with Explicit Error Budgets

Autoregressive transformer inference retains key and value tensors from earlier tokens so each new token can attend to prior context without recomputing those projections. As context length and concurrent sequence count rise, this KV cache can become a substantial part of accelerator memory. KV cache quantization stores those tensors at reduced precision and reconstructs approximations when attention consumes them. The memory arithmetic is attractive, but the resulting error is not a generic model-weight perturbation. Quantized keys affect attention scores before the softmax, while quantized values affect the weighted sum after attention probabilities have been formed.

Artificial Intelligence 15 Sep 2026 4 min read

Preserve Attention Sinks in Streaming Transformers

A bounded attention cache creates a specific failure mode in autoregressive transformers: removing every old token can disturb attention even when those tokens no longer carry useful task content. Some early positions can absorb attention mass across many later queries. If cache eviction removes them, generation quality can degrade more than their semantic value would suggest. These positions are often called attention sinks. The practical implication is narrow but useful: a streaming cache can keep a small prefix of sink positions while rotating the rest of its capacity through recent tokens.

Artificial Intelligence 15 Sep 2026 6 min read

Measure Anisotropy in Embedding Spaces

Embedding vectors can occupy a narrow cone instead of spreading evenly across their available dimensions. In that geometry, unrelated items may still have noticeably positive cosine similarity because many vectors share a common directional component. This concentration is called anisotropy. For developers, anisotropy matters at the point where vector geometry becomes an application signal. A similarity threshold, nearest-neighbor ranking, clustering rule, or novelty detector inherits the distribution produced by the embedding model. The same cosine value can carry different meaning across representation spaces with different directional concentration.

Artificial Intelligence 15 Sep 2026 5 min read

Isolate Attention Across Packed Training Sequences

Padding can consume a large share of a training batch when sequence lengths vary. Sequence packing reduces that waste by placing multiple short samples into one token block. The arithmetic is attractive: more non-padding tokens fit into the same fixed-length tensor. Packing also changes the structure seen by attention. A standard causal mask only prevents a token from attending to later positions. It does not know that two adjacent spans came from separate samples. Without an additional boundary constraint, a token in the second span can attend to tokens from the first span.

Artificial Intelligence 15 Sep 2026 6 min read

Control Length Bias in Beam Search Scoring

Beam search compares multiple partial outputs while autoregressive generation advances token by token. A common scoring rule adds token log probabilities along each candidate sequence. That rule is mathematically consistent with sequence probability, but it also creates a structural preference that developers can miss: extending a sequence normally makes its accumulated log score smaller. This matters whenever candidates of different lengths compete. A decoder can rank a short completed sequence above a longer candidate even when the longer candidate is more useful for the application. Length normalization and length penalties modify that ranking, but they also change the objective being optimized.

Artificial Intelligence 15 Sep 2026 6 min read

Clip Gradient Norms With Clear Scope

Gradient norm clipping changes an optimizer update only when the measured gradient norm exceeds a chosen threshold. The operation sounds local, but its behavior depends on a broader implementation choice: which gradients participate in the norm. Two training loops can use the same threshold and optimizer yet produce different updates because they clip different parameter groups or clip at different points in the update cycle. That makes clipping scope part of the optimization definition, not just a guard against unusually large gradients.

Artificial Intelligence 15 Sep 2026 6 min read

Account for Label Smoothing in Classifier Confidence

A classifier trained with one-hot targets is rewarded for moving probability mass toward the labeled class. Cross-entropy keeps decreasing as the model assigns that class a probability closer to one, even after the predicted class is already correct. Label smoothing changes this pressure by assigning a small amount of target mass to the other classes. That change is easy to treat as a minor detail in the loss function. It is not minor when an application consumes the model’s probability values. The smoothed target changes the optimum encouraged by the training objective, so confidence scores from a smoothed model should not be interpreted as if they came from the same objective as ordinary one-hot training.

Artificial Intelligence 15 Sep 2026 5 min read

Account for Hubness in Embedding Retrieval

An embedding index can return the same few items for many unrelated queries. Their similarity scores may look ordinary, and the nearest-neighbor algorithm may be operating correctly. The distortion can come from the geometry of the representation itself: some vectors become neighbors of unusually many other vectors. This effect is commonly called hubness. Hubness matters because nearest-neighbor retrieval is usually interpreted locally. A query asks which stored vectors are closest to it, but the index does not normally expose how often each candidate also appears near other queries. A candidate that repeatedly occupies neighbor lists can receive more retrieval opportunities than its semantic relevance warrants.

Artificial Intelligence 14 Sep 2026 7 min read

Trade Activation Memory for Recomputation with Checkpointing

Backpropagation needs intermediate values from the forward pass to compute parameter and input gradients. Keeping every required activation alive until its gradient is calculated can consume substantial device memory, especially as model depth, batch size, or sequence length grows. Activation checkpointing changes which intermediates are retained. Selected boundary tensors remain available, while activations inside a checkpointed region are discarded after the forward pass and produced again when the backward pass reaches that region. The model computes the same conceptual function, but the execution schedule exchanges additional computation for lower activation storage.

Artificial Intelligence 14 Sep 2026 6 min read

Track Model Weights with an Exponential Moving Average

Optimizer updates can move model parameters back and forth even when the broader trajectory changes more gradually. An exponential moving average, or EMA, keeps a second parameter state that follows those updates with smoothing. The trainable model still receives ordinary optimizer updates; the EMA state is a derived copy used separately, often for evaluation or export. The mechanism is compact, but its behavior depends on decay, update frequency, initialization, and which state is actually saved or evaluated.

Artificial Intelligence 14 Sep 2026 5 min read

Tie Input Embeddings to the Output Projection

A language model can contain two large matrices indexed by the same vocabulary: one maps token IDs into embedding vectors, while another maps hidden states into vocabulary logits. Weight tying makes those roles share parameters instead of maintaining two independent matrices. The change is compact in code, but it affects parameter counting, gradient flow, dimensional constraints, checkpoint handling, and any component that assumes the input and output weights are separate.

Artificial Intelligence 14 Sep 2026 6 min read

Speculative Decoding Trades Draft Accuracy for Target Model Work

Autoregressive generation normally commits tokens one position at a time. Even when a large model has ample parallel compute available, each next-token decision depends on the prefix produced so far. That serial dependency makes decoding latency sensitive to the number of target-model passes. Speculative decoding changes the unit of work. A cheaper draft model proposes several candidate tokens, then the target model evaluates the proposed continuation in one pass. Accepted candidates advance generation by multiple positions without requiring one separate target pass per accepted token.

Artificial Intelligence 14 Sep 2026 7 min read

Reduce Repetition with Unlikelihood Training

An autoregressive language model is usually trained to increase the probability of the observed next token. That positive objective does not directly state which plausible but unwanted alternatives should receive less probability. When repetitive tokens or phrases remain locally probable, ordinary next-token training can leave generation with a strong route back into content that has already appeared. Unlikelihood training adds a negative signal for selected candidates. Instead of only rewarding the target token, the objective can also penalize tokens chosen because they represent an unwanted behavior, such as repetition within the generated prefix.

Artificial Intelligence 14 Sep 2026 6 min read

Reduce KV Cache Memory with Grouped-Query Attention

Autoregressive transformer inference stores key and value vectors from earlier tokens so each new token does not have to recompute them. With standard multi-head attention, every attention head has its own key and value projections, so the KV cache grows with the number of key-value heads. Grouped-query attention changes that head layout. It keeps multiple query heads but lets several query heads share one key head and one value head. The result reduces cached key-value state without collapsing all query heads into a single shared projection.

Artificial Intelligence 14 Sep 2026 5 min read

Prevent Cross-Example Attention in Packed Sequences

Short training examples can waste much of a fixed-length transformer batch on padding. Sequence packing reduces that waste by placing several examples into one token buffer, but concatenation alone changes the computation. A causal mask prevents a token from attending to future positions; it does not prevent that token from attending to an earlier, unrelated example. The distinction matters whenever packed examples are intended to remain independent. The token buffer may be contiguous for storage and compute while attention, position handling, and loss accounting still need explicit example boundaries.

Artificial Intelligence 14 Sep 2026 5 min read

Normalize Hidden States with RMSNorm

A hidden-state vector can grow or shrink in magnitude as it passes through a neural network. RMSNorm controls that scale by dividing the vector by its root mean square magnitude, then applying a trainable gain. Unlike LayerNorm, it does not subtract the vector mean before rescaling. That missing centering operation is the defining distinction. RMSNorm constrains scale while leaving a uniform shift across coordinates present in the normalized representation. RMSNorm uses the second raw moment For a hidden vector x with d coordinates, its root mean square is:

Artificial Intelligence 14 Sep 2026 6 min read

Measure Embedding Anisotropy Before Trusting Cosine Similarity

Cosine similarity is often treated as if a score has the same meaning across any embedding space. That assumption breaks when vectors occupy a narrow region of the available geometry. If many embeddings share a strong common direction, unrelated items can receive positive cosine scores simply because both align with that direction. This behavior is usually described as embedding anisotropy. It is not a defect in cosine similarity itself. The issue is that cosine measures angles in the representation it receives, including global structure that may have little value for the downstream comparison.

Artificial Intelligence 14 Sep 2026 6 min read

Encode Token Distance with Rotary Position Embeddings

Transformer attention has no intrinsic notion that one token sits three positions before another. Rotary position embeddings, usually called RoPE, inject position into attention by rotating pairs of query and key coordinates before their dot product is computed. The mechanism is easy to reduce to a helper function, yet several details determine its actual behavior: queries and keys must use compatible rotations, each coordinate pair has its own angular frequency, offsets emerge through the dot product, and changing the position scale changes the geometry seen by attention.

Artificial Intelligence 14 Sep 2026 6 min read

Detect Feature Outliers with Mahalanobis Distance

A feature vector can sit close to a reference mean in Euclidean distance and still be unusual for the distribution that produced the reference data. Mahalanobis distance accounts for this by scaling displacement according to covariance. Directions with little observed variation contribute more to the score than directions in which the reference data naturally spreads out. That behavior makes the distance useful as a compact outlier score for model features or embeddings, provided the reference statistics are meaningful and the covariance estimate is numerically usable.

Artificial Intelligence 14 Sep 2026 7 min read

Control Target Certainty with Label Smoothing

A classifier trained with ordinary cross-entropy often receives a one-hot target: probability mass 1 on the labeled class and 0 on every other class. That target keeps rewarding movement toward a more extreme prediction even after the correct class already has the highest score. Label smoothing changes the target distribution rather than the model architecture. A small amount of target mass is moved away from the labeled class and assigned to other classes. Cross-entropy then optimizes against this softened distribution, so the gradient no longer treats absolute certainty on the labeled class as the target state.

Artificial Intelligence 14 Sep 2026 5 min read

Control Classifier Logits with Cosine Normalization

A linear classification head mixes two signals in each logit: the angle between a feature vector and a class weight vector, and the magnitudes of both vectors. Cosine normalization removes the magnitude terms, so class scores depend on directional alignment instead. That change is small in code but substantial in interpretation. Feature norm no longer increases every class comparison merely by growing, class-weight norm no longer acts as an implicit class-specific scale, and the overall sharpness of the softmax must be supplied separately.

Artificial Intelligence 14 Sep 2026 6 min read

Control Beam Search Length Bias with Score Normalization

Beam search usually ranks partial sequences by accumulated token log probability. That score has a built-in dependence on sequence length: each additional token contributes another log probability that is zero or negative. As a result, raw cumulative scores can favor shorter completed sequences even when a longer candidate is preferable for the application. Length normalization changes the ranking rule rather than the model distribution. The distinction matters because decoding can produce different outputs without changing a single model parameter or next-token probability.