Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 13 Sep 2026 6 min read

Clip Gradient Norms Before Optimizer Updates

A training update can become dominated by a gradient whose magnitude is far larger than the range seen in nearby iterations. Global norm clipping places a bound on that update signal before the optimizer consumes it. The operation is simple, but its behavior depends on what is included in the norm, where clipping occurs, and how it interacts with gradient accumulation and mixed-precision scaling. Norm clipping does not repair the source of unstable gradients. It changes the vector passed to the optimizer when its norm exceeds a chosen threshold.

Artificial Intelligence 13 Sep 2026 5 min read

Calibrate Classifier Confidence with Temperature Scaling

A classifier can choose the correct class yet attach a probability that is too concentrated or too diffuse for the application using that score. Temperature scaling addresses this mismatch after model training by applying one positive scalar to the logits before softmax. The mechanism is deliberately narrow. It changes probability sharpness, not the information represented by the classifier. That boundary makes temperature scaling useful when class ranking is acceptable but confidence values need separate calibration.

Artificial Intelligence 13 Sep 2026 7 min read

Build Classification Sets with Split Conformal Prediction

A classifier normally returns one label or a vector of class scores. Neither output directly states how many labels should remain plausible when the system needs a controlled error rate. Split conformal prediction adds a calibration layer that turns those scores into prediction sets. The useful property is not that every individual set has a fixed probability of containing the correct label. Under the standard exchangeability assumption, split conformal methods target marginal coverage across new examples. That distinction shapes both implementation and interpretation.

Artificial Intelligence 13 Sep 2026 8 min read

Balance Sparse Mixture-of-Experts Routing Under Capacity Limits

A sparse mixture-of-experts layer does not send every token through every parameter block. A router scores the available experts, selects a small subset for each token, and dispatches token representations only to those selected experts. That conditional computation is the main attraction of sparse MoE designs, but it also creates a resource-allocation problem inside the model. The router can prefer the same experts for many tokens. Hardware, meanwhile, has finite buffers and communication capacity. A routing policy that looks reasonable from token scores alone can therefore create overloaded experts, idle experts, uneven communication, or discarded assignments.

Artificial Intelligence 12 Sep 2026 11 min read

Understand Sparse Mixture-of-Experts Routing

Understand Sparse Mixture-of-Experts Routing A larger neural network can represent more functions, but using every parameter for every input makes each forward pass expensive. Sparse mixture-of-experts (MoE) models take a different approach: keep many parameter groups available, then activate only a small subset for each token. That sounds like a simple efficiency trick, but routing changes much more than arithmetic cost. It affects training stability, accelerator communication, memory requirements, batching, and the meaning of a model’s total parameter count.

Artificial Intelligence 12 Sep 2026 8 min read

Test Transformer Circuits with Activation Patching

A transformer can expose a clear internal pattern without that pattern being responsible for the output under inspection. Activation patching addresses this gap by changing an internal state and measuring the downstream effect. Instead of asking whether a feature is visible at a layer, it asks whether replacing a selected state changes a defined model behavior. The method is simple in form but sensitive to experimental design. A patch has meaning only relative to the paired inputs, the patched location, the replacement value, and the output metric. Changing any of those can change the causal question being tested.

Artificial Intelligence 12 Sep 2026 9 min read

Stream Long LLM Sessions with Attention Sinks

Stream Long LLM Sessions with Attention Sinks Long-running LLM sessions create a simple resource problem: every generated token can add keys and values to the attention cache. Keep the entire history and memory use keeps growing. Keep only the newest tokens and some transformer models degrade sharply once older cache entries disappear. Attention sinks provide a useful middle ground for compatible models. Instead of retaining the full KV cache, keep a small group of initial tokens plus a moving window of recent tokens. The cache stays bounded, yet the model can remain much more stable than with a recent-token window alone.

Artificial Intelligence 12 Sep 2026 9 min read

Stabilize Image Model Predictions with Test-Time Augmentation

Stabilize Image Model Predictions with Test-Time Augmentation An image classifier can give slightly different answers when the same subject is cropped, mirrored, or resized in a way that preserves its meaning. If those transformations are valid for the task, relying on one view leaves useful evidence unused. Test-time augmentation (TTA) runs inference on several valid views of one input and combines their predictions into a final result. TTA is simple to describe, but safe use depends on details that are easy to miss. A transformation must preserve the target, structured outputs may need to be mapped back before aggregation, probability averaging can affect calibration, and every extra view consumes inference capacity.

Artificial Intelligence 12 Sep 2026 7 min read

Speculative Decoding Depends on Draft Acceptance

Autoregressive generation normally commits one token after each model pass, creating a serial dependency across the output sequence. Speculative decoding changes that execution pattern. A cheaper draft process proposes several future tokens, then the target model evaluates those proposals together and determines which tokens can be committed. The attraction is fewer serial target-model iterations per generated token. That does not make speculative decoding an automatic latency reduction. Its useful operating point depends on how cheaply candidates are produced, how many survive verification, and how much extra work the target model performs while checking them.

Artificial Intelligence 12 Sep 2026 7 min read

Soften Classification Targets with Label Smoothing

A classifier trained with one-hot targets receives a strong signal to push the target class probability toward one and every other class probability toward zero. Cross-entropy supports that behavior even after the predicted class is already correct: making the target probability more extreme can still reduce the loss. Label smoothing changes the target distribution before cross-entropy is computed. Instead of assigning all target mass to one class, it reserves a small amount for the remaining classes. This alters the gradient applied to the logits and reduces pressure toward extreme output distributions.

Artificial Intelligence 12 Sep 2026 10 min read

Route Uncertain Classifier Predictions with Selective Classification

Route Uncertain Classifier Predictions with Selective Classification A classifier does not have to answer every request. In systems where a bad prediction is costly, forcing a label on every input can be a poor product decision even when the model has strong average accuracy. Selective classification adds a reject option. The system returns a model prediction only when an acceptance rule considers the case suitable; otherwise it abstains and sends the case to a fallback such as human review, a second model, or a request for more information.

Artificial Intelligence 12 Sep 2026 8 min read

Reduce Neural Network Inference Cost with Early Exits

Reduce Neural Network Inference Cost with Early Exits A deep neural network normally applies every block to every input, even when an intermediate representation already supports a confident prediction. Early-exit inference changes that fixed-depth path. It attaches prediction heads at intermediate points and lets selected inputs stop before the final block. The appeal is conditional computation: easy cases can consume less compute while ambiguous cases retain access to the full network. The difficult part is deciding when an intermediate prediction is reliable enough to return. A poor exit policy can save computation by silently moving errors toward the shallow heads.

Artificial Intelligence 12 Sep 2026 6 min read

Pack Transformer Training Sequences Without Cross-Sample Attention

Transformer batches often waste token slots on padding when examples have uneven lengths. Sequence packing reduces that waste by placing several shorter samples into one fixed-length token block. The arithmetic is attractive, but concatenation alone changes the training problem: tokens from one sample can attend to tokens from another unless the packed representation preserves sample boundaries. A correct packing scheme therefore has two jobs. It must fill token capacity more densely, and it must keep the model’s effective computation consistent with the intended independence of the original samples.

Artificial Intelligence 12 Sep 2026 7 min read

Inspect Transformer Layer Predictions with the Logit Lens

A decoder-only transformer normally exposes token logits only after its final block and output normalization. The residual stream inside earlier blocks has the same model-width shape, which makes another operation possible: take an intermediate state, apply the model’s output-side normalization when required, and project that state through the output matrix. The resulting vocabulary scores form the logit lens. They provide a token-space view of an internal representation before the remaining transformer blocks have processed it. That view is useful for inspecting how candidate tokens change across depth, but it is not a record of tokens that the model has secretly selected in advance.

Artificial Intelligence 12 Sep 2026 10 min read

Improve LLM Generation with Contrastive Decoding

Improve LLM Generation with Contrastive Decoding A language model can assign high probability to text that is fluent but bland, repetitive, or overly driven by common patterns. Sampling adds variety, but increasing randomness can also admit weak continuations. Contrastive decoding takes a different route: compare a stronger model with a weaker reference model at each generation step, then favor tokens that the stronger model supports more distinctly. The method changes decoding rather than model parameters. It can therefore be useful when you control inference for compatible models and want to experiment with generation quality without another training run. The extra model pass is not free, and the method needs a plausibility guard to avoid promoting strange tokens.

Artificial Intelligence 12 Sep 2026 7 min read

Hubness Can Distort Nearest-Neighbor Embedding Retrieval

Embedding retrieval usually treats each query independently: encode the query, compare it with stored vectors, then return the closest items. That local view can miss a collection-level pattern. Some stored vectors may appear in the nearest-neighbor lists of many unrelated queries far more often than other vectors. This pattern is called hubness. A hub is not necessarily a broadly relevant item. It is a vector that becomes a neighbor unusually often under the representation and distance geometry in use. For developers, the distinction matters because a retrieval pipeline can compute cosine similarity or Euclidean distance exactly as specified and still produce systematically repetitive candidates.

Artificial Intelligence 12 Sep 2026 7 min read

Extend RoPE Context with Position Interpolation

A transformer that uses rotary position embeddings can accept a larger token buffer at the serving layer and still behave poorly at positions far beyond the range used during model training. The tensor shapes may be valid while the positional phases presented to attention are outside the regime the model adapted to. Position interpolation addresses that mismatch by compressing a longer sequence’s position indices into the original position interval before applying RoPE. It does not add memory to the architecture, and it does not make long-context behavior equivalent to native training at the extended length. It changes the positional coordinates supplied to attention.

Artificial Intelligence 12 Sep 2026 9 min read

Extend RoPE Context Windows with Position Interpolation

Extend RoPE Context Windows with Position Interpolation A RoPE-based language model trained on sequences up to a fixed length can behave poorly when inference suddenly asks it to process much larger position indices. The tokens are valid, but the positional pattern can move far outside the range used during training. Position interpolation changes that geometry. Instead of sending larger position indices directly into rotary position embeddings, it compresses a longer sequence into the positional range the model already uses. With suitable adaptation, this can extend the usable context window without changing the transformer architecture.

Artificial Intelligence 12 Sep 2026 6 min read

Exit Transformer Classifiers Early with Entropy Thresholds

A transformer classifier normally sends every input through every layer, even when an intermediate representation already supports a concentrated class prediction. Entropy-based early exit changes that fixed-depth behavior. Prediction heads attached to intermediate layers estimate class distributions, and inference can stop once a distribution passes a configured entropy threshold. The mechanism makes model depth input-dependent. Some inputs may leave after relatively few layers, while uncertain inputs continue through more of the network. That flexibility also introduces a new source of error: an intermediate head can be confident and still be wrong.

Artificial Intelligence 12 Sep 2026 7 min read

Embedding Anisotropy Can Compress Cosine Score Separation

Two embedding vectors can have a high cosine similarity even when the items they represent are not close in the task-specific sense a retrieval system needs. One source of this mismatch is embedding anisotropy: vectors are distributed unevenly across representation space, often with substantial mass concentrated around shared directions. Cosine similarity removes vector magnitude from the comparison, but it does not remove a common directional component. If many vectors point partly in the same direction, unrelated pairs can start from an elevated cosine baseline. The useful distinction between relevant and irrelevant items then has to appear within a narrower score range.

Artificial Intelligence 12 Sep 2026 9 min read

Control LLM Behavior with Activation Steering

Control LLM Behavior with Activation Steering Prompting controls a language model through its input tokens. Fine-tuning changes model parameters. Activation steering offers a third option: change selected internal activations while the model runs, without rewriting its weights. That makes activation steering useful for experiments where you want to test whether an internal direction is connected to a behavior, or apply a lightweight behavior shift during generation. It also creates new engineering questions. A steering vector can help at one layer and damage output at another. A strength that works on short prompts can become excessive on different inputs. A behavioral shift can also come with losses in fluency or task accuracy.

Artificial Intelligence 12 Sep 2026 7 min read

Control Expert Routing in Mixture-of-Experts Models

A mixture-of-experts layer can contain far more parameters than it evaluates for each token. A router scores a set of experts, selects a small subset, and sends each token only to those selected computation paths. The parameter count can grow without making every token execute every expert. That sparse structure creates a separate systems problem: the router decides where computation lands. Two models with the same experts and the same nominal top-k routing can have very different behavior if one spreads tokens across experts and the other concentrates them on a few paths.

Artificial Intelligence 12 Sep 2026 6 min read

Control Diffusion Conditioning with Classifier-Free Guidance

A conditional diffusion model can follow its conditioning signal more strongly at sampling time without a separate classifier. Classifier-free guidance does this by evaluating a model in conditional and unconditional modes, then amplifying the difference between those predictions. That difference is the central mechanism. The guidance scale does not simply make a prompt louder in an abstract sense. It changes the denoising prediction along a direction defined by what the conditioning input contributes relative to an unconditional prediction at the same noisy state.

Artificial Intelligence 12 Sep 2026 6 min read

Contrast Language Model Logits with Expert-Amateur Decoding

A language model can assign high probability to tokens that are fluent but generic. Contrastive decoding changes token selection by comparing a stronger expert model with a weaker amateur model at the same generation position. A token becomes attractive when the expert favors it more strongly than the amateur does. The comparison is not an unrestricted subtraction across the vocabulary. The original method also keeps candidate tokens inside a plausibility set defined by the expert. That constraint matters because a large expert-amateur score gap can otherwise promote a token that both models consider implausible.