Skip to content

Archive

Model Architecture

7 articles
Artificial Intelligence 24 Sep 2026 5 min read

SwiGLU Gates Transformer Feed-Forward Channels with a Second Projection

SwiGLU splits a transformer feed-forward input into two projected paths, applies SiLU to one path, then multiplies the two results element by element. The second projection is not an auxiliary statistic: its values directly gate the activated path before the output projection. For hidden state x, a common structural form is: g = SiLU(x W_gate) u = x W_up h = g * u y = h W_down Bias terms, projection orientation, intermediate width, and parameter names vary across architectures. The defining boundary is the element-wise product between a nonlinear projected branch and another projected branch.

Artificial Intelligence 24 Sep 2026 5 min read

Sliding-Window Attention Bounds Each Query to a Local Token Horizon

A causal attention layer does not always expose every earlier token to every query. With a sliding window of width w, the query at position i can be restricted to recent positions rather than the full prefix. The attention graph becomes local: old tokens fall outside the direct edge set even though they remain part of the sequence. That boundary changes computation, memory traffic, and information paths at the same time. It is not merely an optimized implementation of full attention. Once the mask removes distant key-value pairs, the layer implements a different dependency pattern.

Artificial Intelligence 24 Sep 2026 5 min read

RMSNorm Scales Hidden States Without Mean Centering

RMSNorm normalizes a hidden-state vector without subtracting its coordinate mean. That single omission separates it from LayerNorm at the mathematical interface: RMSNorm controls scale through a root-mean-square statistic, while any common offset across coordinates remains part of the transformed representation. For a vector x with width D, a common RMSNorm form is: rms(x) = sqrt(mean(x_i^2) + eps) y_i = g_i * x_i / rms(x) Here g_i is a trainable per-coordinate scale and eps is a small positive term defined by the model implementation. Exact parameterization and numeric details belong to the checkpoint and runtime contract.

Artificial Intelligence 24 Sep 2026 5 min read

RMSNorm Rescales Hidden States Without Mean Centering

RMSNorm rescales a vector from its root mean square without first subtracting the vector mean. That missing centering operation is the defining difference from LayerNorm: both can control vector scale, but only LayerNorm explicitly shifts the normalized coordinates around a zero sample mean. For a hidden vector x with width d, a common RMSNorm form is: rms = sqrt((1/d) * sum(x_i^2) + epsilon) y_i = gain_i * x_i / rms The exact placement of epsilon, numeric precision used for the reduction, and presence of extra affine terms depend on the implementation. The structural operation remains division by an RMS statistic rather than division by a standard deviation computed after mean subtraction.

Artificial Intelligence 24 Sep 2026 5 min read

MoE Expert Capacity Bounds Token Routing

A sparse Mixture-of-Experts layer can contain many expert networks while activating only a small subset for each token. That conditional computation depends on a router, but router scores alone do not determine the executed graph. In implementations with bounded expert batches, each expert also has a finite number of token slots. This creates a second boundary after expert selection: a token can prefer an expert that has no remaining capacity. The handling of that overflow is an implementation and architecture choice with direct consequences for training and serving.

Artificial Intelligence 08 Sep 2026 8 min read

Residual Connections in Deep Neural Networks

Making a neural network deeper gives it more transformations to work with, but depth alone does not make optimization easy. A stack of layers must learn useful transformations while gradients travel backward through every stage. As the stack grows, that optimization path can become difficult even when the deeper model has enough capacity to represent a good solution. Residual connections change what a block is asked to learn. Instead of making the block produce an entirely new representation, they let it learn a change to the representation it already received. The original input travels along a shortcut and is added back to the learned branch.

Artificial Intelligence 06 Sep 2026 8 min read

Reduce Language Model Parameters with Weight Tying

Language models need to turn token IDs into vectors before processing them and turn hidden vectors back into vocabulary scores before predicting the next token. A straightforward design gives those two operations separate parameter matrices. When the vocabulary and hidden dimension are large, each matrix can contain many parameters. Weight tying removes that duplication by reusing one parameter matrix for both roles. The input side reads rows from the matrix as token embeddings; the output side uses the same learned vectors to score candidate tokens, usually through the matrix transpose.