Skip to content

Archive

Neural Networks

96 articles
Artificial Intelligence 12 Sep 2026 7 min read

Soften Classification Targets with Label Smoothing

A classifier trained with one-hot targets receives a strong signal to push the target class probability toward one and every other class probability toward zero. Cross-entropy supports that behavior even after the predicted class is already correct: making the target probability more extreme can still reduce the loss. Label smoothing changes the target distribution before cross-entropy is computed. Instead of assigning all target mass to one class, it reserves a small amount for the remaining classes. This alters the gradient applied to the logits and reduces pressure toward extreme output distributions.

Artificial Intelligence 12 Sep 2026 8 min read

Reduce Neural Network Inference Cost with Early Exits

Reduce Neural Network Inference Cost with Early Exits A deep neural network normally applies every block to every input, even when an intermediate representation already supports a confident prediction. Early-exit inference changes that fixed-depth path. It attaches prediction heads at intermediate points and lets selected inputs stop before the final block. The appeal is conditional computation: easy cases can consume less compute while ambiguous cases retain access to the full network. The difficult part is deciding when an intermediate prediction is reliable enough to return. A poor exit policy can save computation by silently moving errors toward the shallow heads.

Artificial Intelligence 12 Sep 2026 7 min read

Control Expert Routing in Mixture-of-Experts Models

A mixture-of-experts layer can contain far more parameters than it evaluates for each token. A router scores a set of experts, selects a small subset, and sends each token only to those selected computation paths. The parameter count can grow without making every token execute every expert. That sparse structure creates a separate systems problem: the router decides where computation lands. Two models with the same experts and the same nominal top-k routing can have very different behavior if one spreads tokens across experts and the other concentrates them on a few paths.

Artificial Intelligence 12 Sep 2026 10 min read

Compress Neural Networks with Knowledge Distillation

Compress Neural Networks with Knowledge Distillation A model can meet your quality target in a notebook and still be too expensive to serve. A large network may consume too much memory, add unacceptable latency, or make high request volume costly. Knowledge distillation addresses this problem by using a stronger model, called the teacher, to guide the training of a smaller student model. The key idea is richer than copying the teacher’s final answer. The teacher produces a distribution across possible outputs, and that distribution can reveal useful relationships between alternatives. A student can train against those soft targets while also using the original labels.

Artificial Intelligence 12 Sep 2026 7 min read

Calibrate Neural Classifier Confidence with Temperature Scaling

A neural classifier can choose the correct class often enough for an application while assigning probabilities that are too concentrated or too diffuse. Accuracy alone does not expose this mismatch. A system that acts differently at confidence thresholds also depends on the numerical probabilities attached to its predictions. Temperature scaling is a post-training calibration method that adjusts the sharpness of classifier logits with one positive scalar. For a fixed input, it preserves the ordering of logits, so the predicted class remains unchanged when ordinary argmax decoding is used. What changes is the probability distribution produced after softmax.

Artificial Intelligence 11 Sep 2026 10 min read

Understand Wide Neural Networks with the Neural Tangent Kernel

Understand Wide Neural Networks with the Neural Tangent Kernel A neural network may contain millions of parameters, yet a useful theoretical view asks a smaller question: when one training example changes the parameters, how does that update affect the prediction for another example? The neural tangent kernel (NTK) answers that question through gradients. It measures how similarly two inputs respond to an infinitesimal parameter update. In a particular infinite-width regime, this kernel becomes effectively fixed during training, turning a nonlinear parameter-optimization problem into a much simpler kernel process in function space.

Artificial Intelligence 11 Sep 2026 10 min read

Train Through Discrete Decisions with the Straight-Through Estimator

Train Through Discrete Decisions with the Straight-Through Estimator Neural networks are usually trained by following gradients through a chain of differentiable operations. A hard discrete choice breaks that chain. Rounding a value, selecting a binary gate, or quantizing an activation can make the forward computation useful while leaving ordinary backpropagation with a zero or undefined derivative at the decision. The straight-through estimator (STE) is a practical workaround. It keeps the discrete operation in the forward pass but substitutes a simpler derivative during the backward pass. That makes optimization possible, at the cost of using a gradient that is not the true derivative of the forward computation.

Artificial Intelligence 11 Sep 2026 9 min read

Score Candidates with Energy-Based Models

Score Candidates with Energy-Based Models Many AI systems need to decide which candidate fits an input: which reply matches a conversation, which label fits an image, or which configuration is plausible. A common design makes the model output a probability directly. Energy-based models take a more general route: they assign each input-candidate pair a scalar energy, with lower values representing greater compatibility. That simple change is useful because the model can focus on relative preference without requiring every architecture to produce a normalized probability during scoring. It also introduces real engineering challenges. Training needs informative alternatives, probability normalization can be expensive, and inference may require searching a large candidate space.

Artificial Intelligence 11 Sep 2026 10 min read

RMSNorm in Transformers

RMSNorm in Transformers A transformer repeatedly adds residual updates to its hidden states. Without some way to control the scale of those values, training deep networks becomes harder to manage. Normalization layers are one of the mechanisms used to keep that computation well behaved. RMSNorm, short for root mean square normalization, is a normalization method used in many transformer architectures. It looks similar to LayerNorm, but it deliberately leaves out one operation: subtracting the mean. Instead, RMSNorm measures the root mean square magnitude of a hidden vector and rescales the vector by that magnitude.

Artificial Intelligence 11 Sep 2026 9 min read

Focus Classifier Training with Focal Loss

Focus Classifier Training with Focal Loss A classifier can spend much of its training signal on examples it already handles correctly. This becomes especially troublesome when easy examples vastly outnumber difficult ones. A detector, for instance, may encounter many obvious background locations for every location containing an object. Focal loss changes the contribution of each example according to the model’s confidence in the correct class. Easy, high-confidence examples receive less weight. Harder examples retain more of their cross-entropy loss. The mechanism is small, but using it well requires understanding what it changes and what it doesn’t.

Artificial Intelligence 11 Sep 2026 10 min read

Cut Neural Network Inference Cost with Early Exits

Cut Neural Network Inference Cost with Early Exits A neural network usually spends the same depth of computation on every input. That is convenient, but not every input needs the same effort. A clear image of a stop sign may be classified correctly after relatively shallow processing, while an occluded sign may need the full network. Early-exit inference adds intermediate prediction points to a model and lets sufficiently confident inputs stop before the final layer. The aim is not to make every request cheaper. It is to spend less computation on easier cases while preserving a deeper path for harder ones.

Artificial Intelligence 11 Sep 2026 9 min read

Control Neural Network Weight Scale with Spectral Normalization

Control Neural Network Weight Scale with Spectral Normalization A neural network layer can amplify a small change in its input into a much larger change in its output. Large amplification isn’t automatically a defect, but it can make some models harder to control during training, especially when one network is reacting to another as in a generative adversarial network. Spectral normalization puts a direct constraint on that amplification for a linear transformation. It rescales a weight matrix using its largest singular value, called the spectral norm. The result is a simple mechanism with a precise local meaning: under the Euclidean norm, the normalized linear map cannot stretch a vector by more than the chosen scale.

Artificial Intelligence 11 Sep 2026 11 min read

Clip Gradients Relative to Parameter Scale with AGC

Clip Gradients Relative to Parameter Scale with AGC A gradient can be large in absolute terms without being large for the parameter it updates. A gradient norm of 0.1 is modest next to a parameter norm of 10, but enormous next to a parameter norm of 0.001. Ordinary gradient clipping doesn’t see that distinction: it compares gradients with a fixed threshold. Adaptive gradient clipping (AGC) uses a different reference point. It compares a gradient’s norm with the norm of the parameter unit that gradient will update. If the gradient is too large relative to the parameter, AGC rescales it before the optimizer step.

Artificial Intelligence 10 Sep 2026 9 min read

Regularize Neural Networks with Mixup

Regularize Neural Networks with Mixup A neural network can fit its training examples while behaving unpredictably in the space between them. If two nearby inputs belong to different classes, standard training tells the model what to do at the endpoints but often says little about intermediate points. Mixup changes that training signal. Instead of training only on individual examples, it creates synthetic examples by interpolating pairs of inputs and their labels. The model is then asked to make a correspondingly mixed prediction. This acts as a regularizer because it constrains how predictions may change between training examples.

Artificial Intelligence 10 Sep 2026 9 min read

Prepare Models for Low-Precision Inference with Quantization-Aware Training

A model can work well in floating point and lose useful accuracy after its weights or activations are quantized for deployment. The problem is not mysterious: rounding and clipping change the numbers that flow through the network, while the original model was optimized without those changes in the loop. Quantization-aware training (QAT) exposes the model to an approximation of those low-precision numerics while its parameters can still adapt. Training remains differentiable in floating point, but the forward computation simulates the quantization errors expected after conversion.

Artificial Intelligence 10 Sep 2026 10 min read

Neural Collapse in Deep Classifiers

Neural Collapse in Deep Classifiers A classifier can keep changing after it already predicts every training example correctly. Cross-entropy loss can continue to fall, feature vectors can reorganize, and the final classification layer can become increasingly regular. Looking only at training accuracy hides all of that movement. Neural collapse is a name for a collection of geometric patterns that can emerge late in the training of deep classifiers. The striking part isn’t simply that examples from the same class become similar. Under the conditions where neural collapse appears, within-class variation can shrink while class centers and classifier weights approach a highly symmetric arrangement.

Artificial Intelligence 09 Sep 2026 12 min read

Train Discrete Neural Operations with Straight-Through Estimators

Neural networks are usually trained with gradient descent, which depends on small changes in parameters producing informative changes in the loss. A discrete operation can break that assumption. Rounding a value, choosing a binary gate, or selecting a quantized level may be exactly what the forward computation needs, yet its derivative can be zero almost everywhere or undefined at transition points. A straight-through estimator (STE) is a practical way to keep training in that situation. The forward pass uses the discrete operation, while the backward pass substitutes a simpler derivative so that a gradient can flow through it. The important consequence is easy to miss: the backward signal is generally not the true derivative of the discrete forward computation. It is a deliberately chosen surrogate.

Artificial Intelligence 09 Sep 2026 9 min read

Diagnose and Prevent Dying ReLU Units

ReLU is one of the simplest neural-network activation functions: negative inputs become zero and positive inputs pass through unchanged. That simplicity makes optimization efficient, but it creates a failure mode that can quietly waste model capacity. A unit can move into a state where its pre-activation is negative for every relevant training example, so its ReLU output stays zero and the unit stops receiving a useful gradient through that activation.

Artificial Intelligence 08 Sep 2026 11 min read

Vanishing and Exploding Gradients in Deep Networks

A deep neural network can have enough capacity to solve a task and still fail to learn because useful training signals do not reach all of its layers. Parameters near the output may update normally while earlier layers receive gradients that are almost zero. In the opposite case, gradients can grow so large that one optimizer step destabilizes the model. These are the vanishing-gradient and exploding-gradient problems. They are not simply labels for “training is bad.” They describe what happens to derivatives as backpropagation repeatedly applies the chain rule through many transformations.

Artificial Intelligence 08 Sep 2026 9 min read

Use Test-Time Augmentation for More Stable Predictions

A classifier can change its prediction because an object moved a few pixels, an image was cropped differently, or another harmless transformation changed the input representation. If those transformations should not change the correct answer, that sensitivity is undesirable. Test-time augmentation (TTA) addresses this problem by running the same trained model on several meaning-preserving versions of an input and combining their predictions. Instead of asking the model for one view of the evidence, TTA asks it to evaluate several valid views.

Artificial Intelligence 08 Sep 2026 10 min read

Train Neural Networks with Curriculum Learning

Most training pipelines treat the dataset as a fixed pool and repeatedly shuffle it. That is a strong default: it is simple, exposes the model to the full data distribution, and avoids assumptions about which examples should come first. But some learning problems have a useful notion of progression. A model may learn basic patterns more reliably before it is asked to handle noisy, ambiguous, or structurally difficult examples. Curriculum learning makes that progression explicit. Instead of changing the model architecture or the loss, it changes which training examples are emphasized at different stages of training. A common curriculum begins with easier examples and gradually introduces harder ones.

Artificial Intelligence 08 Sep 2026 11 min read

Trade Compute for Memory with Activation Checkpointing

A neural network can fit comfortably in accelerator memory for inference and still run out of memory during training. The reason is that training needs more than model weights. Backpropagation also needs intermediate values from the forward pass, and those activations can consume a large share of memory in deep models or with long sequences and large batches. Activation checkpointing reduces that memory pressure by deliberately not keeping every intermediate activation. Instead, training saves selected checkpoints and recomputes missing forward-pass values when the backward pass needs them. The trade is straightforward: keep fewer activations in memory, but perform extra computation.

Artificial Intelligence 08 Sep 2026 9 min read

The Softmax Bottleneck in Language Models

A language model can have a powerful network behind it and still be constrained by the layer that turns its hidden state into next-token probabilities. In the common linear-softmax output layer, that constraint has a precise form: across many contexts, the model can represent only a limited family of log-probability patterns. This limitation is known as the softmax bottleneck. It is not a claim that softmax itself is defective, nor does it mean every modern language model is visibly harmed by it. It is a structural result about a particular output parameterization.

Artificial Intelligence 08 Sep 2026 9 min read

Stabilize Neural Network Weights with Exponential Moving Averages

A neural network’s parameters rarely move smoothly toward their final values. Mini-batch training produces noisy updates: one batch may push a weight in one direction, while the next pushes it partly back. The final checkpoint therefore represents one point on a noisy training path, not necessarily the most useful point near the end of that path. An exponential moving average (EMA) of model weights keeps a second set of parameters that changes more slowly than the actively trained model. Recent training states contribute more than old ones, but no single update immediately replaces the averaged weights.