Clip Gradient Norms Before Optimizer Updates
A training update can become dominated by a gradient whose magnitude is far larger than the range seen in nearby iterations. Global norm clipping places a bound on that update signal before the optimizer consumes it. The operation is simple, but its behavior depends on what is included in the norm, where clipping occurs, and how it interacts with gradient accumulation and mixed-precision scaling. Norm clipping does not repair the source of unstable gradients. It changes the vector passed to the optimizer when its norm exceeds a chosen threshold.