Skip to content

Archive

Softmax

8 articles
Artificial Intelligence 24 Sep 2026 5 min read

Temperature Scaling Recalibrates Classifier Confidence Without Changing Class Order

A classifier can rank the correct class above every alternative yet attach probabilities that are systematically too concentrated or too diffuse. Temperature scaling addresses that mismatch after training by applying one scalar to the logits before softmax. It changes reported confidence without changing the underlying classifier parameters. The mechanism is narrow. It does not repair incorrect class rankings, add information to the representation, or make every individual probability accurate. Its target is the relationship between confidence and observed outcomes on data representative of deployment.

Artificial Intelligence 24 Sep 2026 6 min read

Tanh Logit Soft Capping Bounds Extreme Scores Before Softmax

A softmax can accept logits of any finite magnitude, but large score gaps make its output increasingly concentrated. Tanh logit soft capping inserts a bounded nonlinear transform before softmax so that no transformed logit exceeds a configured magnitude. For a positive cap c, a common form is: softcap(z; c) = c * tanh(z / c) The operation does not clip at a hard threshold. It behaves almost linearly near zero and gradually compresses larger magnitudes as they approach -c or c.

Artificial Intelligence 24 Sep 2026 4 min read

QK Normalization Bounds Attention Logit Scale Before Softmax

QK normalization inserts normalization on query and key vectors before the attention dot product. The operation changes the geometry of the score calculation: vector magnitude no longer enters the dot product in the same unrestricted form, while directional alignment remains part of the score. For one query vector q and key vector k, ordinary scaled dot-product attention forms a score such as: s = dot(q, k) / sqrt(D) A QK-normalized variant first applies the model’s specified normalization functions:

Artificial Intelligence 23 Sep 2026 5 min read

Temperature Scaling Changes Softmax Sharpness Without Changing Logit Order

A decoder can assign the same ranking to every token before and after temperature scaling while producing substantially different probabilities. The mechanism is simple: for positive temperature T, logits are divided by T before softmax. Division by the same positive scalar preserves order, but softmax converts the changed gaps between logits into a different probability distribution. That distinction matters in inference systems because temperature does not select tokens by itself. Its visible effect depends on what happens after the scaled softmax: direct sampling, top-k filtering, top-p filtering, greedy selection, or another decoding rule.

Artificial Intelligence 23 Sep 2026 4 min read

Temperature Scaling Changes Sampling Without Changing Logit Order

Temperature is often exposed as a single generation parameter, but its effect is narrower than a general control for output quality. For a fixed vector of finite logits, positive temperature rescales the gaps before softmax. It changes the resulting probabilities without changing which logit is larger than another. That distinction matters when a serving layer combines temperature with greedy selection, top-k filtering, top-p filtering, penalties, or implementation-specific handling of zero temperature. The same numeric setting can participate in a different decoding pipeline even though the underlying scaling operation is simple.

Artificial Intelligence 23 Sep 2026 3 min read

Stable Softmax Subtracts the Maximum Logit Before Exponentiation

A decoder can receive logits large enough that direct exponentiation is numerically unsafe even though the intended probability distribution is ordinary. Softmax does not require exponentiating the original values. Subtracting the largest logit from every logit produces the same distribution in exact real arithmetic while moving the exponentials into a safer numeric range. This shift is a property of softmax itself, not a model-specific heuristic. A shared shift cancels during normalization For logits (z_1,\ldots,z_n), softmax assigns

Artificial Intelligence 23 Sep 2026 6 min read

Attention Logit Scaling Keeps Dot Products in a Stable Softmax Range

A dot product between a query and a key tends to grow in magnitude as their dimension grows. In scaled dot-product attention, the score is divided by the square root of the key dimension before softmax. That factor is not a cosmetic normalization. It controls the scale presented to softmax under a specific statistical assumption about the query and key components. The familiar expression is Attention(Q, K, V) = softmax(QK^T / sqrt(d_k)) V where d_k is the query-key dimension for one attention head. The scaling term affects the distribution of attention probabilities even though it does not change the ordering of logits by itself.

Artificial Intelligence 23 Sep 2026 6 min read

Additive Logit Bias Changes Token Odds Before Sampling

A decoder can favor or suppress a token without changing model weights. Add a constant to that token’s logit before softmax, and its probability changes relative to the rest of the vocabulary. The operation is simple, but its effect depends on where the bias enters the decoding pipeline and on every transformation that follows it. This makes additive logit bias useful as an inference control, but not as a general semantic constraint. It changes a score used by the decoder. It does not rewrite the model’s internal representation or guarantee that a concept disappears from generated text.