Skip to content

Archive

Deep Learning

2 articles
Artificial Intelligence 03 Sep 2026 10 min read

Train Larger AI Models with Gradient Accumulation

Training a neural network often becomes memory-bound before it becomes compute-bound. You may want a batch of 64 examples for stable optimization, but the model, activations, optimizer state, and input tensors leave enough accelerator memory for only 8 examples at a time. Reducing the batch size to 8 may work, but it also changes the optimization process. Gradient accumulation provides another option: process several smaller microbatches, add their gradients together, and update the model only after the desired effective batch has been processed.

Artificial Intelligence 03 Sep 2026 8 min read

Activation Functions in Transformer Feed-Forward Networks

Attention gets much of the attention in transformer explanations, but every transformer layer also contains a feed-forward network that performs substantial computation on each token representation. The activation function inside that network is a small-looking design choice with an important job: it introduces nonlinearity so the network can learn transformations that stacked linear projections alone cannot express. The feed-forward block matters when reading model architectures, comparing implementations, estimating parameter and compute costs, or deciding whether two designs are actually equivalent.