Skip to content

Archive

Neural Networks

96 articles
Artificial Intelligence 08 Sep 2026 8 min read

Residual Connections in Deep Neural Networks

Making a neural network deeper gives it more transformations to work with, but depth alone does not make optimization easy. A stack of layers must learn useful transformations while gradients travel backward through every stage. As the stack grows, that optimization path can become difficult even when the deeper model has enough capacity to represent a good solution. Residual connections change what a block is asked to learn. Instead of making the block produce an entirely new representation, they let it learn a change to the representation it already received. The original input travels along a shortcut and is added back to the learned branch.

Artificial Intelligence 08 Sep 2026 9 min read

Regularize Residual Networks with Stochastic Depth

Deep residual networks can overfit even when their skip connections make optimization manageable. Standard dropout can regularize individual activations, but residual architectures offer another useful unit to randomize: the entire residual branch. Stochastic depth randomly removes selected residual branches during training while keeping the skip path intact. A training example may therefore pass through a slightly shallower effective network on one step and the full set of blocks on another. At inference time, every residual branch is normally active.

Artificial Intelligence 08 Sep 2026 10 min read

Model Soups for Combining Fine-Tuned Models

A hyperparameter sweep often leaves you with several fine-tuned models that are individually useful. The usual workflow keeps the checkpoint with the best validation score and discards the rest. An ensemble can use several checkpoints, but then every request may require multiple model evaluations, increasing inference cost and operational complexity. A model soup offers a third option: average the parameters of compatible fine-tuned models and deploy the resulting parameter set as one model. The technique is simple, but its simplicity can be misleading. Parameter averaging is meaningful only when the checkpoints are sufficiently compatible, and the averaged model still needs independent evaluation.

Artificial Intelligence 08 Sep 2026 9 min read

Masked Autoencoders for Visual Representation Learning

Labeled image datasets are expensive to build, but unlabeled images are often plentiful. A useful pretraining strategy is therefore to create a learning signal from each image itself instead of asking a human to annotate it. A masked autoencoder (MAE) does this by hiding part of an image and training a model to reconstruct the missing content. The reconstruction task is not usually the final product. Its purpose is to make the encoder learn visual representations that can later support tasks such as classification or detection.

Artificial Intelligence 08 Sep 2026 10 min read

Dynamical Isometry in Neural Networks

A deep neural network can look well scaled one layer at a time and still be difficult to optimize. Signals pass through many transformations, and small expansions or contractions can multiply with depth. By the time a gradient travels through the whole network, some directions may have nearly disappeared while others have been amplified dramatically. Dynamical isometry gives a precise way to reason about this problem. Instead of asking only whether the average gradient magnitude is reasonable, it asks how the network transforms different directions in its input space. The relevant object is the input-output Jacobian, and its singular values provide the main measurements.

Artificial Intelligence 08 Sep 2026 8 min read

Deep Ensembles for Model Uncertainty

A neural network can return a confident prediction even when an input is unfamiliar or ambiguous. Looking only at one model’s largest probability can therefore hide an important question: would another plausible model, trained on the same task, make the same decision? A deep ensemble helps answer that question by training several neural networks independently and combining their predictions. The combined prediction can improve robustness in some settings, while disagreement among members provides a practical uncertainty signal. It is not a guarantee that the prediction is correct, and it does not detect every kind of uncertainty.

Artificial Intelligence 07 Sep 2026 12 min read

Xavier and He Initialization for Neural Networks

A deep neural network can fail before learning has had a fair chance. If its initial weights make activations or gradients shrink layer after layer, useful signals can become tiny. If those quantities grow too much, training can become unstable. The optimizer may receive a problem that is unnecessarily difficult even though the architecture and data are otherwise reasonable. Weight initialization tries to start the network in a numerically useful regime. Two common schemes are Xavier initialization, also called Glorot initialization, and He initialization, also called Kaiming initialization. Both choose the scale of random weights from the size of a layer, but they make different assumptions about how signals pass through the activation function.

Artificial Intelligence 07 Sep 2026 8 min read

Weight Tying in Language Models

A language model needs to solve two related problems with vocabulary-sized parameters. At the input, it must turn each token ID into a vector. At the output, it must turn a hidden state into one score for every possible next token. A straightforward architecture gives these two operations separate parameter matrices. That works, but it can be expensive when the vocabulary and hidden dimension are large. Weight tying is a simple architectural idea: use the same learned matrix for the input token embeddings and the output token projection when their shapes and semantics are compatible.

Artificial Intelligence 07 Sep 2026 9 min read

Weight Decay in Adam and AdamW

A training configuration can contain a parameter named weight_decay without making it obvious what operation the optimizer actually performs. That ambiguity matters most with adaptive optimizers such as Adam: adding an L2 penalty to the loss and directly decaying parameters are not generally the same update. The distinction is easy to miss because the two procedures are closely related under ordinary stochastic gradient descent (SGD). Once an optimizer rescales different coordinates using gradient history, however, the equivalence breaks.

Artificial Intelligence 07 Sep 2026 9 min read

Transfer Hyperparameters Across Model Width with MuP

Scaling a neural network creates an expensive tuning problem. A learning rate that works for a small prototype may behave differently after hidden dimensions become much wider. If every model size needs a fresh hyperparameter sweep, experimenting on small models saves less compute than it first appears. Maximal Update Parametrization, usually written MuP or μP, addresses this problem by changing how parameter initialization and learning rates scale with model width. The goal is not to make a wider model identical to a narrow one. It is to make important training dynamics behave consistently enough that hyperparameters tuned on a smaller proxy can often transfer to a wider target.

Artificial Intelligence 07 Sep 2026 10 min read

Train Neural Networks with Sharpness-Aware Minimization

A neural network can reach low training loss at parameter values where a small change to the weights makes the loss rise sharply. Standard optimization does not directly discourage this behavior: it mainly asks whether the loss is low at the current parameters. Sharpness-Aware Minimization (SAM) changes the training objective. Instead of optimizing only the loss at the current weights, it approximately optimizes the worst loss in a small neighborhood around them. The practical idea is simple: find a nearby parameter perturbation that makes the current mini-batch harder, then update the original model using the gradient measured at those perturbed parameters.

Artificial Intelligence 07 Sep 2026 10 min read

Resolve Conflicting Gradients in Multi-Task Learning

Training one neural network to solve several tasks can reduce duplicated computation and let related tasks share useful representations. It also creates a problem that single-task training does not have: two losses can ask the same shared parameter to move in opposing directions during the same update. Simply adding the losses does not make that disagreement disappear. Their gradients are added too, so one task can partially cancel another or dominate the shared update. Gradient surgery is a family of techniques that changes task gradients before combining them. A well-known example is projected conflicting gradients, commonly called PCGrad, which removes a conflicting component of one task’s gradient relative to another.

Artificial Intelligence 07 Sep 2026 10 min read

Probe Neural Network Representations with Linear Classifiers

A neural network can produce the right output while leaving an important engineering question unanswered: what information exists inside its intermediate representations? Suppose an image classifier predicts product categories. You may want to know whether an early layer already separates shapes, whether a later layer distinguishes categories, or whether a supposedly irrelevant attribute such as camera source remains easy to recover. Looking only at the final prediction does not answer those questions.

Artificial Intelligence 07 Sep 2026 9 min read

Prevent Catastrophic Forgetting in Continual Learning

Updating a neural network with new data sounds straightforward: continue training on the new examples and deploy the improved model. The difficulty is that an update which helps the new data can damage behavior the model learned earlier. A classifier that learns a new group of products, for example, may become worse at recognizing older groups even though those old classes never changed. This failure is called catastrophic forgetting. It matters most in continual learning, where a model learns from a sequence of tasks or data distributions instead of training once on a fixed mixed dataset.

Artificial Intelligence 07 Sep 2026 10 min read

Neural Network Pruning: Sparsity, Structure, and Speed

A neural network can contain parameters that contribute little to its useful predictions. Removing some of them can reduce storage or computation, but there is an important trap: a model with fewer nonzero weights is not automatically a model that runs faster. That distinction matters when developers use pruning to compress a trained network. The pruning rule determines what disappears, the hardware and runtime determine whether the resulting structure can be exploited, and the evaluation procedure determines whether the saved resources are worth any quality loss.

Artificial Intelligence 07 Sep 2026 9 min read

Knowledge Distillation for Smaller Classifiers

A model can be accurate enough for a product and still be too expensive to deploy. A large classifier may exceed a mobile memory budget, miss a latency target, or cost too much when every request requires substantial compute. Replacing it with a smaller model reduces those costs, but training the smaller model only from ground-truth labels can leave useful information behind. Knowledge distillation addresses this problem by training a smaller student model to learn from a stronger teacher model. Instead of seeing only the correct class, the student can also learn how the teacher distributes its confidence across the alternatives.

Artificial Intelligence 07 Sep 2026 12 min read

Gradient Noise Scale and Batch Size

Increasing a neural network’s batch size can make more accelerators useful, but the benefit does not grow indefinitely. At some point, processing more examples before each update gives a cleaner estimate of nearly the same gradient direction while consuming additional examples and compute. Gradient noise scale provides a useful mental model for this transition. It compares the variation in per-example gradients with the strength of their average. When gradient estimates are noisy relative to their mean, averaging more examples can remove meaningful noise. When they are already stable, a larger batch has less statistical work left to do.

Artificial Intelligence 07 Sep 2026 9 min read

Gradient Centralization in Neural Network Training

Neural network optimizers normally consume the gradients produced by backpropagation directly. Gradient centralization inserts one small transformation between those two steps: for selected weight tensors, it subtracts the mean of each gradient vector before the optimizer uses it. That operation is easy to implement, but its effect is easy to misunderstand. It is not gradient clipping, because it does not cap large values. It is not normalization, because it does not divide by a norm or standard deviation. It changes the direction of the update by removing one particular component.

Artificial Intelligence 07 Sep 2026 10 min read

Diagnose and Prevent Dead ReLU Neurons

ReLU is a simple and effective activation function, but its simplicity creates a failure mode that is easy to miss. A neuron can reach a state where its pre-activation is negative for every relevant input. Its ReLU output is then always zero, and the gradient through that activation is also zero. If this persists, the neuron may stop participating in learning. This is commonly called a dead ReLU or dying ReLU problem. It does not mean that every zero activation is a defect: sparse activations are a normal consequence of ReLU. The useful question is whether a unit is inactive only for some inputs or effectively inactive across the data it needs to model.

Artificial Intelligence 07 Sep 2026 11 min read

Compress Neural Network Layers with Low-Rank Factorization

Large neural networks spend much of their memory and computation multiplying activations by weight matrices. Some of those matrices contain more independent structure than the model actually needs for a particular deployment. If so, we can approximate one large matrix with two smaller matrices and reduce the number of stored parameters and multiply-add operations. This technique is called low-rank factorization. The central idea is simple, but using it well requires more than choosing a smaller number. Compression changes the weights, approximation error can accumulate through a network, and fewer arithmetic operations do not guarantee lower wall-clock latency on every device.

Artificial Intelligence 07 Sep 2026 10 min read

Average Neural Network Weights with Stochastic Weight Averaging

A neural network rarely finishes training at the only useful point in parameter space. Late in training, stochastic gradient descent (SGD) can visit several nearby parameter settings that all perform reasonably well, while the final checkpoint represents only one of them. Stochastic weight averaging (SWA) turns that observation into a simple training technique: collect model parameters from multiple points late in an SGD trajectory and compute their arithmetic mean. The result is one model with averaged weights, so inference does not require running an ensemble of all collected checkpoints.

Artificial Intelligence 07 Sep 2026 10 min read

Adapt Neural Networks with Gradient Reversal

A classifier can perform well in evaluation and then weaken after deployment because the inputs changed. Product photos may come from a new camera, support messages may use different vocabulary, or sensor readings may come from different hardware. Collecting labels for the new environment can be expensive even when unlabeled examples are easy to obtain. Domain-adversarial training addresses one version of this problem. It asks a feature extractor to support the prediction task while making the source and target domains difficult to distinguish. A gradient reversal layer makes those two goals trainable with ordinary backpropagation by reversing the domain classifier’s gradient before it reaches the feature extractor.

Artificial Intelligence 06 Sep 2026 10 min read

Stabilize Neural Network Evaluation with Exponential Moving Average Weights

A neural network’s final training step is not necessarily its most useful checkpoint. Stochastic optimization keeps moving the parameters as it follows noisy mini-batch gradients, so two nearby checkpoints can behave slightly differently even when training is otherwise healthy. An exponential moving average (EMA) of model weights gives you a second set of parameters that changes more smoothly. Instead of evaluating only the latest training weights, you maintain a weighted history in which recent weights matter most and older weights gradually fade away.

Artificial Intelligence 06 Sep 2026 8 min read

Reduce Language Model Parameters with Weight Tying

Language models need to turn token IDs into vectors before processing them and turn hidden vectors back into vocabulary scores before predicting the next token. A straightforward design gives those two operations separate parameter matrices. When the vocabulary and hidden dimension are large, each matrix can contain many parameters. Weight tying removes that duplication by reusing one parameter matrix for both roles. The input side reads rows from the matrix as token embeddings; the output side uses the same learned vectors to score candidate tokens, usually through the matrix transpose.