Skip to content

Topic archive

Artificial Intelligence

Artificial intelligence includes machine learning, deep learning, and systems that apply models to practical tasks.

396 articles
Artificial Intelligence 12 Sep 2026 10 min read

Compress Neural Networks with Knowledge Distillation

Compress Neural Networks with Knowledge Distillation A model can meet your quality target in a notebook and still be too expensive to serve. A large network may consume too much memory, add unacceptable latency, or make high request volume costly. Knowledge distillation addresses this problem by using a stronger model, called the teacher, to guide the training of a smaller student model. The key idea is richer than copying the teacher’s final answer. The teacher produces a distribution across possible outputs, and that distribution can reveal useful relationships between alternatives. A student can train against those soft targets while also using the original labels.

Artificial Intelligence 12 Sep 2026 7 min read

Calibrate Neural Classifier Confidence with Temperature Scaling

A neural classifier can choose the correct class often enough for an application while assigning probabilities that are too concentrated or too diffuse. Accuracy alone does not expose this mismatch. A system that acts differently at confidence thresholds also depends on the numerical probabilities attached to its predictions. Temperature scaling is a post-training calibration method that adjusts the sharpness of classifier logits with one positive scalar. For a fixed input, it preserves the ordering of logits, so the predicted class remains unchanged when ordinary argmax decoding is used. What changes is the probability distribution produced after softmax.

Artificial Intelligence 12 Sep 2026 7 min read

Bound KV Cache Growth in Streaming LLM Inference

Autoregressive transformer inference normally keeps key and value states from earlier tokens so each new token can attend to prior context without recomputing those states. The cache grows with sequence length. For a long-running stream, that growth eventually becomes a memory constraint even when generation itself continues one token at a time. KV cache eviction puts a bound on that state by discarding selected cached positions. The memory effect is straightforward: fewer retained key-value pairs occupy less cache space. The model effect is more subtle. Once a position is removed, later attention layers cannot use its cached key and value in the ordinary attention calculation. An eviction policy therefore changes both resource use and the effective attention history.

Artificial Intelligence 12 Sep 2026 10 min read

Align LLMs with Direct Preference Optimization

Align LLMs with Direct Preference Optimization Supervised fine-tuning works well when you can provide a target response for each prompt. It becomes less natural when the signal is comparative: one answer is preferred over another, but neither is a perfect target to copy. Direct Preference Optimization (DPO) turns those preference pairs into a training objective for a language model without requiring a separately trained reward model or an online reinforcement step.

Artificial Intelligence 11 Sep 2026 9 min read

Use Token Dropout to Train Robust Sequence Models

Use Token Dropout to Train Robust Sequence Models A sequence model can become too dependent on a few easy input clues. Remove one field, truncate a message, or corrupt a token at inference time, and a prediction that looked reliable on clean validation data may change sharply. Token dropout is a simple training-time corruption technique: randomly hide some input tokens while keeping the learning target unchanged. The model is forced to solve some training examples without every usual clue. Used carefully, this can reduce brittle dependence on individual tokens. Used carelessly, it can destroy information the task genuinely requires.

Artificial Intelligence 11 Sep 2026 10 min read

Understand Wide Neural Networks with the Neural Tangent Kernel

Understand Wide Neural Networks with the Neural Tangent Kernel A neural network may contain millions of parameters, yet a useful theoretical view asks a smaller question: when one training example changes the parameters, how does that update affect the prediction for another example? The neural tangent kernel (NTK) answers that question through gradients. It measures how similarly two inputs respond to an infinitesimal parameter update. In a particular infinite-width regime, this kernel becomes effectively fixed during training, turning a nonlinear parameter-optimization problem into a much simpler kernel process in function space.

Artificial Intelligence 11 Sep 2026 10 min read

Train Through Discrete Decisions with the Straight-Through Estimator

Train Through Discrete Decisions with the Straight-Through Estimator Neural networks are usually trained by following gradients through a chain of differentiable operations. A hard discrete choice breaks that chain. Rounding a value, selecting a binary gate, or quantizing an activation can make the forward computation useful while leaving ordinary backpropagation with a zero or undefined derivative at the decision. The straight-through estimator (STE) is a practical workaround. It keeps the discrete operation in the forward pass but substitutes a simpler derivative during the backward pass. That makes optimization possible, at the cost of using a gradient that is not the true derivative of the forward computation.

Artificial Intelligence 11 Sep 2026 9 min read

Score Candidates with Energy-Based Models

Score Candidates with Energy-Based Models Many AI systems need to decide which candidate fits an input: which reply matches a conversation, which label fits an image, or which configuration is plausible. A common design makes the model output a probability directly. Energy-based models take a more general route: they assign each input-candidate pair a scalar energy, with lower values representing greater compatibility. That simple change is useful because the model can focus on relative preference without requiring every architecture to produce a normalized probability during scoring. It also introduces real engineering challenges. Training needs informative alternatives, probability normalization can be expensive, and inference may require searching a large candidate space.

Artificial Intelligence 11 Sep 2026 10 min read

RMSNorm in Transformers

RMSNorm in Transformers A transformer repeatedly adds residual updates to its hidden states. Without some way to control the scale of those values, training deep networks becomes harder to manage. Normalization layers are one of the mechanisms used to keep that computation well behaved. RMSNorm, short for root mean square normalization, is a normalization method used in many transformer architectures. It looks similar to LayerNorm, but it deliberately leaves out one operation: subtracting the mean. Instead, RMSNorm measures the root mean square magnitude of a hidden vector and rescales the vector by that magnitude.

Artificial Intelligence 11 Sep 2026 11 min read

Reweight Long-Tailed Classification with Effective Sample Counts

A classifier trained on a long-tailed dataset can see thousands of examples from common classes and only a handful from rare ones. Ordinary empirical risk minimization gives the common classes more influence simply because they appear more often. A tempting fix is to weight each class by the inverse of its example count, but that can make a tiny class disproportionately influential, including any mislabeled examples it contains. Class-balanced loss based on the effective number of samples provides a smoother way to derive class weights. Instead of treating every additional example as equally informative, it models diminishing returns within a class and weights classes according to an adjusted, or effective, sample count.

Artificial Intelligence 11 Sep 2026 9 min read

Reduce Vision Transformer Compute with Token Merging

Reduce Vision Transformer Compute with Token Merging Vision transformers can spend substantial computation processing many patch tokens that carry similar information. A patch covering one part of a clear sky may produce a representation close to nearby sky patches, yet ordinary self-attention continues to process each token separately. Token merging reduces that redundancy by combining selected tokens as they move through the network. Unlike token pruning, which removes tokens, merging tries to preserve their information in a smaller set of representations. The practical goal is simple: reduce the token count in later transformer blocks while keeping task quality within an acceptable range.

Artificial Intelligence 11 Sep 2026 11 min read

Prevent VAE Posterior Collapse with Free Bits

Prevent VAE Posterior Collapse with Free Bits A variational autoencoder can appear to train normally while its latent representation becomes nearly useless. The decoder learns to explain the data without depending on the latent variable, the encoder moves toward the prior, and the KL divergence shrinks toward zero. This failure mode is called posterior collapse. Free bits is a small change to the VAE objective that can reduce one source of that collapse. It stops the KL term from rewarding the optimizer for squeezing an already-small amount of latent information even closer to zero. The technique is simple, but its name and common shorthand can lead to a misleading mental model. Free bits does not force a latent variable to contain a chosen amount of information. It changes the optimization pressure below a threshold.

Artificial Intelligence 11 Sep 2026 11 min read

Let Classifiers Abstain with Selective Prediction

Let Classifiers Abstain with Selective Prediction A classifier does not have to answer every case. In many systems, forcing a prediction on an ambiguous input is worse than sending that input to a human, requesting more information, or falling back to a safer workflow. Selective prediction gives a model that option. The classifier produces its usual prediction, but the system accepts it only when a selection rule considers the case reliable enough. Otherwise, the system abstains.

Artificial Intelligence 11 Sep 2026 10 min read

Fuse Keyword and Vector Search with Reciprocal Rank Fusion

Fuse Keyword and Vector Search with Reciprocal Rank Fusion A RAG system often needs two kinds of retrieval at once. Keyword search is good at exact strings such as product codes, error messages, and names. Vector search can recover passages that express the same idea with different words. Running both is easy; combining their scores correctly is where many implementations become fragile. Reciprocal rank fusion (RRF) solves that problem by ignoring raw scores and combining rank positions instead. BM25 and vector similarity do not share a stable numeric scale. The sections below calculate RRF on a small example, then cover the parameters and evaluation checks that matter in hybrid retrieval.

Artificial Intelligence 11 Sep 2026 9 min read

Focus Classifier Training with Focal Loss

Focus Classifier Training with Focal Loss A classifier can spend much of its training signal on examples it already handles correctly. This becomes especially troublesome when easy examples vastly outnumber difficult ones. A detector, for instance, may encounter many obvious background locations for every location containing an object. Focal loss changes the contribution of each example according to the model’s confidence in the correct class. Easy, high-confidence examples receive less weight. Harder examples retain more of their cross-entropy loss. The mechanism is small, but using it well requires understanding what it changes and what it doesn’t.

Artificial Intelligence 11 Sep 2026 10 min read

Cut Neural Network Inference Cost with Early Exits

Cut Neural Network Inference Cost with Early Exits A neural network usually spends the same depth of computation on every input. That is convenient, but not every input needs the same effort. A clear image of a stop sign may be classified correctly after relatively shallow processing, while an occluded sign may need the full network. Early-exit inference adds intermediate prediction points to a model and lets sufficiently confident inputs stop before the final layer. The aim is not to make every request cheaper. It is to spend less computation on easier cases while preserving a deeper path for harder ones.

Artificial Intelligence 11 Sep 2026 9 min read

Control Neural Network Weight Scale with Spectral Normalization

Control Neural Network Weight Scale with Spectral Normalization A neural network layer can amplify a small change in its input into a much larger change in its output. Large amplification isn’t automatically a defect, but it can make some models harder to control during training, especially when one network is reacting to another as in a generative adversarial network. Spectral normalization puts a direct constraint on that amplification for a linear transformation. It rescales a weight matrix using its largest singular value, called the spectral norm. The result is a simple mechanism with a precise local meaning: under the Euclidean norm, the normalized linear map cannot stretch a vector by more than the chosen scale.

Artificial Intelligence 11 Sep 2026 11 min read

Clip Gradients Relative to Parameter Scale with AGC

Clip Gradients Relative to Parameter Scale with AGC A gradient can be large in absolute terms without being large for the parameter it updates. A gradient norm of 0.1 is modest next to a parameter norm of 10, but enormous next to a parameter norm of 0.001. Ordinary gradient clipping doesn’t see that distinction: it compares gradients with a fixed threshold. Adaptive gradient clipping (AGC) uses a different reference point. It compares a gradient’s norm with the norm of the parameter unit that gradient will update. If the gradient is too large relative to the parameter, AGC rescales it before the optimizer step.

Artificial Intelligence 11 Sep 2026 10 min read

Build Prediction Sets with Conformal Prediction

Build Prediction Sets with Conformal Prediction A classifier usually returns one label even when several labels are plausible. That is convenient for software interfaces, but it can hide uncertainty exactly where mistakes are expensive. A document router might be unsure between billing and account, yet an argmax still emits one of them. Conformal prediction offers another interface: return a set of labels sized according to the evidence. Easy inputs can produce one label. Ambiguous inputs can produce several. Under specific assumptions, the procedure also gives a finite-sample coverage guarantee.

Artificial Intelligence 11 Sep 2026 9 min read

Adapt Language Models with Prefix Tuning

Adapt Language Models with Prefix Tuning Full fine-tuning changes a model’s weights for each task. That can be effective, but storing and serving a separate full checkpoint for every task becomes expensive as model size and task count grow. Prefix tuning offers a different arrangement: keep the pretrained model frozen and train a small set of task-specific states that participate in attention. This article builds a practical mental model for prefix tuning, shows how it differs from text prompts and low-rank weight adapters, and explains the trade-offs that matter when training or serving several task variants.

Artificial Intelligence 10 Sep 2026 10 min read

Whiten Embeddings Without Breaking Vector Search

Whiten Embeddings Without Breaking Vector Search Embedding search can produce a vector space whose dimensions are strongly correlated or whose variance is concentrated in a few directions. When that geometry interferes with retrieval, embedding whitening is one possible post-processing step: center the vectors, rotate them into uncorrelated directions, and rescale those directions to comparable variance. The transformation is simple to describe but easy to misuse. Whitening changes the geometry that your similarity function sees. If you fit it on the wrong data, transform only one side of retrieval, or keep unstable low-variance directions, search quality can get worse even though the transformed covariance looks cleaner.

Artificial Intelligence 10 Sep 2026 10 min read

Use Late Interaction for Fine-Grained Neural Retrieval

Use Late Interaction for Fine-Grained Neural Retrieval A single text embedding is convenient: encode a query into one vector, encode each document into one vector, then rank documents by vector similarity. That design scales well, but compression happens early. A paragraph containing several distinct ideas must squeeze all of them into one fixed-size representation before the query arrives. Late interaction keeps more of that detail. Instead of representing each text with only one vector, it retains multiple contextual token vectors and compares them at retrieval time. The document can still be encoded ahead of time, but the final relevance score is computed from fine-grained query-to-document matches.

Artificial Intelligence 10 Sep 2026 10 min read

Temperature in LLM Sampling

A language model can produce very different continuations from the same prompt even when its weights and context haven’t changed. One of the controls behind that variation is temperature. Temperature is often described as a creativity knob. That description is convenient but incomplete. Temperature doesn’t add ideas to a model, improve its knowledge, or directly control factual accuracy. It changes the probability distribution used to choose the next token. The practical effect depends on what the model already considers plausible at that step.

Artificial Intelligence 10 Sep 2026 9 min read

Regularize Neural Networks with Mixup

Regularize Neural Networks with Mixup A neural network can fit its training examples while behaving unpredictably in the space between them. If two nearby inputs belong to different classes, standard training tells the model what to do at the endpoints but often says little about intermediate points. Mixup changes that training signal. Instead of training only on individual examples, it creates synthetic examples by interpolating pairs of inputs and their labels. The model is then asked to make a correspondingly mixed prediction. This acts as a regularizer because it constrains how predictions may change between training examples.