Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 24 Sep 2026 5 min read

Temperature Scaling Recalibrates Classifier Confidence Without Changing Class Order

A classifier can rank the correct class above every alternative yet attach probabilities that are systematically too concentrated or too diffuse. Temperature scaling addresses that mismatch after training by applying one scalar to the logits before softmax. It changes reported confidence without changing the underlying classifier parameters. The mechanism is narrow. It does not repair incorrect class rankings, add information to the representation, or make every individual probability accurate. Its target is the relationship between confidence and observed outcomes on data representative of deployment.

Artificial Intelligence 24 Sep 2026 6 min read

Tanh Logit Soft Capping Bounds Extreme Scores Before Softmax

A softmax can accept logits of any finite magnitude, but large score gaps make its output increasingly concentrated. Tanh logit soft capping inserts a bounded nonlinear transform before softmax so that no transformed logit exceeds a configured magnitude. For a positive cap c, a common form is: softcap(z; c) = c * tanh(z / c) The operation does not clip at a hard threshold. It behaves almost linearly near zero and gradually compresses larger magnitudes as they approach -c or c.

Artificial Intelligence 24 Sep 2026 5 min read

SwiGLU Gates Transformer Feed-Forward Channels with a Second Projection

SwiGLU splits a transformer feed-forward input into two projected paths, applies SiLU to one path, then multiplies the two results element by element. The second projection is not an auxiliary statistic: its values directly gate the activated path before the output projection. For hidden state x, a common structural form is: g = SiLU(x W_gate) u = x W_up h = g * u y = h W_down Bias terms, projection orientation, intermediate width, and parameter names vary across architectures. The defining boundary is the element-wise product between a nonlinear projected branch and another projected branch.

Artificial Intelligence 24 Sep 2026 6 min read

Speculative Decoding Verifies Draft Tokens Without Changing the Target Sampling Distribution

Autoregressive decoding normally commits one token after each target-model evaluation. A sequence of K generated tokens therefore creates a serial dependency chain: token t+1 cannot be sampled until token t is fixed and becomes part of the prefix. Speculative decoding changes the amount of useful work obtained from a target-model call without removing that causal dependency. A cheaper draft model first proposes several continuation tokens. The target model then evaluates those candidate positions together. Tokens whose draft probabilities are compatible with the target distribution can be accepted, while the first rejected position is corrected using a residual distribution. The resulting samples follow the target model’s distribution when the acceptance and correction procedure is implemented as specified.

Artificial Intelligence 24 Sep 2026 4 min read

Speculative Decoding Verifies Draft Tokens Without Changing the Target Distribution

Autoregressive generation normally invokes the target model once for every emitted token. Speculative decoding changes that execution pattern: a cheaper draft model proposes several tokens, and the target model evaluates the proposed block in a single verification pass. The speed opportunity comes from doing useful target-model work for multiple positions at once, not from treating draft output as authoritative. Draft tokens are proposals, not final output Let the target model define distribution p and the draft model define distribution q at a given position. The draft samples a candidate token from q. Verification then decides whether that candidate can be retained as a sample consistent with p.

Artificial Intelligence 24 Sep 2026 5 min read

Speculative Decoding Verifies Draft Tokens in Parallel

Autoregressive generation normally commits one token before the next target-model step can be evaluated. Speculative decoding changes that execution pattern. A cheaper proposal distribution produces a short candidate block, while the target model evaluates the proposed positions together. An acceptance rule then determines how much of that block can become output. The mechanism is not ordinary batching. Tokens inside the proposal remain autoregressive, and the target model still defines the intended output distribution when an exact speculative-sampling algorithm is used. The useful change is that one expensive verification pass can account for several output positions when enough proposals are accepted.

Artificial Intelligence 24 Sep 2026 5 min read

Sliding-Window Attention Bounds Each Query to a Local Token Horizon

A causal attention layer does not always expose every earlier token to every query. With a sliding window of width w, the query at position i can be restricted to recent positions rather than the full prefix. The attention graph becomes local: old tokens fall outside the direct edge set even though they remain part of the sequence. That boundary changes computation, memory traffic, and information paths at the same time. It is not merely an optimized implementation of full attention. Once the mask removes distant key-value pairs, the layer implements a different dependency pattern.

Artificial Intelligence 24 Sep 2026 5 min read

Sliding-Window Attention Bounds Direct Token Access

A causal transformer does not have to expose every earlier token to every query position. With sliding-window attention, position i can attend only to a bounded range of preceding keys. Tokens outside that range are absent from that layer’s attention operation, even if they still belong to the model input. That boundary changes more than the attention matrix shape. It separates direct access from information that can reach a position only after being propagated through intermediate hidden states.

Artificial Intelligence 24 Sep 2026 5 min read

Rotary Position Embedding Converts Absolute Indices into Relative Attention Phases

Self-attention can compare token content without assigning an order to token positions. Rotary Position Embedding (RoPE) inserts position into that comparison by rotating paired coordinates of queries and keys. Each token receives an absolute rotation angle, yet the query-key inner product reduces those two absolute angles to their difference. That algebraic cancellation is the central mechanism. RoPE does not add a position vector to the hidden state. It changes the orientation of query and key components before their dot product is evaluated.

Artificial Intelligence 24 Sep 2026 5 min read

RoPE Position Interpolation Compresses Indices Before Rotation

Rotary position embeddings apply a position-dependent rotation to query and key components before their dot product is evaluated. If a model was trained with positions up to a reference length, simply sending much larger position indices into the same rotation rule can place attention in angular regimes that were not represented in that training range. Position interpolation changes the coordinate supplied to RoPE. For an extension factor s > 1, a simplified form maps

Artificial Intelligence 24 Sep 2026 4 min read

RoPE Encodes Relative Offsets Through Rotated Query-Key Phases

Rotary Position Embedding (RoPE) applies position-dependent rotations to query and key coordinates before their attention dot product. The resulting score carries relative position through the phase difference between those rotations rather than through an additive position vector attached to the token representation. This distinction is structural. RoPE starts from absolute indices for each rotation, yet the query-key inner product can be written in terms of the offset between their positions.

Artificial Intelligence 24 Sep 2026 5 min read

RMSNorm Scales Hidden States Without Mean Centering

RMSNorm normalizes a hidden-state vector without subtracting its coordinate mean. That single omission separates it from LayerNorm at the mathematical interface: RMSNorm controls scale through a root-mean-square statistic, while any common offset across coordinates remains part of the transformed representation. For a vector x with width D, a common RMSNorm form is: rms(x) = sqrt(mean(x_i^2) + eps) y_i = g_i * x_i / rms(x) Here g_i is a trainable per-coordinate scale and eps is a small positive term defined by the model implementation. Exact parameterization and numeric details belong to the checkpoint and runtime contract.

Artificial Intelligence 24 Sep 2026 5 min read

RMSNorm Rescales Hidden States Without Mean Centering

RMSNorm rescales a vector from its root mean square without first subtracting the vector mean. That missing centering operation is the defining difference from LayerNorm: both can control vector scale, but only LayerNorm explicitly shifts the normalized coordinates around a zero sample mean. For a hidden vector x with width d, a common RMSNorm form is: rms = sqrt((1/d) * sum(x_i^2) + epsilon) y_i = gain_i * x_i / rms The exact placement of epsilon, numeric precision used for the reduction, and presence of extra affine terms depend on the implementation. The structural operation remains division by an RMS statistic rather than division by a standard deviation computed after mean subtraction.

Artificial Intelligence 24 Sep 2026 4 min read

QK Normalization Bounds Attention Logit Scale Before Softmax

QK normalization inserts normalization on query and key vectors before the attention dot product. The operation changes the geometry of the score calculation: vector magnitude no longer enters the dot product in the same unrestricted form, while directional alignment remains part of the score. For one query vector q and key vector k, ordinary scaled dot-product attention forms a score such as: s = dot(q, k) / sqrt(D) A QK-normalized variant first applies the model’s specified normalization functions:

Artificial Intelligence 24 Sep 2026 5 min read

Prefix KV Caching Reuses Only the Shared Token Prefix

A prefix cache hit ends at the first point where a new request can no longer reuse previously computed state. The reusable object is not a piece of source text in isolation. It is model state produced for an ordered token prefix under execution conditions that make that state compatible with the new request. For transformer inference, that state is commonly the key-value cache created during prefill. Reusing it can remove repeated computation for the shared prefix while leaving the divergent suffix to be processed normally.

Artificial Intelligence 24 Sep 2026 5 min read

Prefix Caching Reuses KV State for Identical Token Prefixes

Autoregressive serving often receives requests that start with the same long token sequence. A system prompt, tool schema, fixed document header, or other repeated context can cause the model to compute the same prefix attention state again for each request. Prefix caching targets that repeated prefill work by retaining compatible key-value state and attaching later requests to it. The reuse boundary is exact tokenized state, not semantic similarity. Two prompts that express the same idea but tokenize differently do not produce an interchangeable prefix cache entry.

Artificial Intelligence 24 Sep 2026 6 min read

Position Interpolation Compresses RoPE Indices into the Original Context Range

A RoPE-based Transformer associates token positions with rotations whose angles depend on the position index and per-dimension frequencies. Feeding a sequence beyond the context range used during training pushes those rotations to position indices the model did not encounter in that regime. Position Interpolation changes that boundary by scaling the extended indices back into the original range before the rotary transformation is applied. For an original context limit L and a target context L' > L, a simplified linear mapping is:

Artificial Intelligence 24 Sep 2026 5 min read

PagedAttention Decouples Logical KV Sequences from Physical Cache Blocks

Autoregressive serving keeps a growing key-value state for every active sequence. If that state must occupy one contiguous physical region sized for a request, allocation becomes coupled to uncertain sequence length: reserving too much wastes capacity, while extending or relocating a growing region complicates memory management. PagedAttention changes that allocation boundary. A sequence is represented as logical KV blocks, while its physical blocks may reside at unrelated locations in the cache pool.

Artificial Intelligence 24 Sep 2026 6 min read

Packed Sequences Need Boundary-Aware Attention Masks

Sequence packing reduces padding by placing several variable-length examples into one token buffer. The storage layout may look like one long sequence, but the examples are still semantically independent. A standard causal mask does not preserve that independence by itself. For a decoder-only Transformer, causal masking blocks attention to future positions. It does not normally block attention to earlier positions that belong to another packed example. If boundaries are ignored, tokens in a later example can attend to keys and values from an earlier one. The model then receives context that the data pipeline intended to keep separate.

Artificial Intelligence 24 Sep 2026 5 min read

Multi-Token Prediction Adds Parallel Future-Token Losses to a Shared Model Trunk

A next-token language model normally applies one predictive objective at position t: the hidden representation at that position is used to score token x_(t+1). Multi-token prediction changes that training boundary. One shared model trunk produces the representation, while multiple output heads predict several subsequent tokens from that position. The mechanism adds supervision at multiple future offsets without requiring a separate transformer trunk for every offset. It is therefore a change to the training objective and prediction heads, not a claim that ordinary autoregressive generation can emit several unchecked tokens as one exact step.

Artificial Intelligence 24 Sep 2026 5 min read

Multi-Head Latent Attention Compresses KV State Before Head-Specific Expansion

Multi-Head Latent Attention changes the state retained across autoregressive decoding. Instead of requiring a full key tensor and value tensor for every cached token and attention head, the architecture can retain a lower-dimensional latent representation and derive head-specific key/value information through projection structure. That distinction matters at the cache boundary. Ordinary multi-head attention commonly stores already projected per-head keys and values. MLA moves part of that representation behind a compact bottleneck, so persistent state and compute no longer have the same shape.

Artificial Intelligence 24 Sep 2026 5 min read

MoE Expert Capacity Bounds Token Routing

A sparse Mixture-of-Experts layer can contain many expert networks while activating only a small subset for each token. That conditional computation depends on a router, but router scores alone do not determine the executed graph. In implementations with bounded expert batches, each expert also has a finite number of token slots. This creates a second boundary after expert selection: a token can prefer an expert that has no remaining capacity. The handling of that overflow is an implementation and architecture choice with direct consequences for training and serving.

Artificial Intelligence 24 Sep 2026 4 min read

Label Smoothing Redistributes Target Probability Across Classes

A classifier trained with one-hot targets assigns all target probability mass to one class. Label smoothing changes that target before cross-entropy is evaluated: some mass is moved away from the designated class and assigned to other classes. The network architecture can remain identical, yet the optimization objective is no longer the same. That distinction matters when interpreting confidence, loss values, and implementation settings. Label smoothing is not a post-processing operation on predicted probabilities. It changes the target distribution used to produce the training signal.

Artificial Intelligence 24 Sep 2026 4 min read

KV-Cache Quantization Reduces Cache Bytes with Separate Accuracy and Kernel Costs

Autoregressive decoding retains key and value tensors from earlier tokens, so KV-cache memory grows with retained context. Quantizing those tensors changes a direct term in that memory footprint: fewer bits are stored for each cached element. The trade is not free capacity. Reduced precision adds representation error and requires a concrete scaling, storage, and kernel strategy. Cache precision is separate from weight precision Model weights and KV state have different lifetimes. Weights are persistent across requests, while KV tensors are generated from each request and grow as its sequence advances. A model can therefore use one numerical format for weights and another for its cache.