A repetition penalty changes the next-token distribution without changing the model parameters or hidden state computation. The model still produces its logits from the current context, but the decoder edits selected scores according to tokens that have already appeared. Sampling then operates on the edited scores rather than directly on the model output.

That distinction matters when reproducing generation behavior. Two systems can run identical model weights on identical token IDs and still emit different continuations because their repetition rules differ in formula, token-history scope, or position in the decoding pipeline.

The penalty sits between model output and sampling

Let the model produce a logit z_i for token i. A decoder can transform that value into z'_i when token i is present in the history considered by the penalty. Softmax, top-k, top-p, or another selection rule then receives z'.

The model itself does not assign z'_i. It assigns z_i; the decoding policy performs the transformation afterward. This makes repetition control an inference-time policy rather than an intrinsic probability property of the model.

A generic pipeline can be represented as:

context -> model -> logits -> repetition transform -> filtering -> sampling

Implementations can order score processors differently. If another processor is nonlinear or removes candidates, changing the order can change the final candidate set. Reproducing a serving stack therefore requires the actual processor sequence, not just matching parameter names.

A multiplicative rule needs sign-aware handling

One common repetition-penalty form uses a positive factor p, usually with p > 1 when repetition is meant to be discouraged. For a token already present in the relevant history, the transformation can be written as:

if z_i > 0:
    z'_i = z_i / p
else:
    z'_i = z_i * p

The sign branch is significant. Dividing every logit by p would move negative logits toward zero, which can increase their relative probability. Multiplying negative values by p instead makes them more negative, while dividing positive values reduces them. Both branches lower the penalized token’s score relative to an unchanged zero reference when p > 1.

This transformation is not equivalent to subtracting a fixed constant. Its absolute effect depends on the original logit magnitude and sign. A token with logit 8 changes more in absolute terms than one with logit 0.5 under the same factor.

The exact formula remains implementation-specific. A parameter called repetition_penalty does not by itself establish sign handling, history scope, or interaction with other processors.

History membership and token frequency are different signals

A presence-style repetition rule can depend only on whether a token occurred. Once a token is present, seeing it five more times need not increase the transformation. A frequency-based rule instead can scale its adjustment with the number of prior occurrences.

These policies encode different state. Presence requires a membership test over the selected history. Frequency requires counts. A multiplicative penalty based on membership also differs from an additive frequency adjustment even when both reduce repeated-token probability in a particular example.

This difference becomes visible when a continuation intentionally repeats syntax or identifiers. If a token occurs once in a prompt and many times in generated output, a presence-only policy can treat both situations identically when its history scope includes both regions. A count-based policy can produce a progressively larger adjustment.

Token identity sets the granularity

Repetition processors normally operate on token IDs, not semantic concepts or rendered words. A visible word can consist of one token in one context and several tokens in another. Whitespace, punctuation, capitalization, and tokenizer vocabulary can also produce distinct IDs for text that appears closely related to a person.

As a result, a token-level penalty does not directly mean “avoid repeating this word.” It means that specific token identities receive score transformations when they satisfy the processor’s history rule.

This boundary is especially relevant for code, structured data, and names. Repeated delimiters, indentation tokens, field names, or subword fragments can be penalized even when repetition is structurally required. The decoder has no semantic exception unless the implementation adds one.

The history window changes the effective policy

The processor must define which previous token IDs count as history. Some systems may inspect the full available sequence; others can restrict the check to generated tokens or a bounded recent region. Those choices produce different score vectors from the same current model logits.

A bounded region also makes the policy time-dependent in a specific sense: once an old token leaves the inspected region, it can stop receiving a repetition adjustment even though it remains in the model’s context. Conversely, a prompt-inclusive policy can penalize a token before the model has generated that token even once.

Context visibility and penalty visibility are therefore separate concepts. A token can remain addressable by attention but fall outside the repetition processor’s inspected region, or it can be included in both.

Penalties reshape odds rather than guarantee novelty

Lowering a repeated token’s score does not guarantee that another token will be selected. The token can remain the highest-scoring candidate after the transformation, especially when its original margin is large. Sampling can also select it with nonzero probability if it remains in the candidate set.

The effect depends on the entire transformed score vector and on subsequent decoding operations. Temperature scaling changes score differences; top-k and top-p can alter membership of the candidate set; deterministic selection can choose the maximum remaining score. A repetition factor cannot be interpreted independently from that pipeline.

For the same reason, increasing the factor does not provide a general semantic guarantee against repeated phrases. The mechanism acts on token scores, while phrase repetition emerges from sequences of token choices and the model state produced by prior tokens.

Serving compatibility requires processor semantics

Generation APIs often expose compact decoding parameters, but compatibility requires more than matching numeric values. The serving implementation must agree on the penalty formula, sign behavior, inspected token region, token counting rule, processor order, and tokenizer IDs.

These details are outside the model checkpoint unless the serving format explicitly records them. A checkpoint can therefore be identical across two deployments while generation differs because decoder state is maintained differently.

The practical boundary is precise: repetition penalties modify inference-time token scores using prior token identity or counts. They can bias generation away from previously observed tokens, but they do not rewrite model logits at their source, erase tokens from context, or impose a semantic ban on repeated text. Any reproducibility claim that omits the processor rule leaves part of the decoding function unspecified.