Sequence packing reduces padding by placing several variable-length examples into one token buffer. The storage layout may look like one long sequence, but the examples are still semantically independent. A standard causal mask does not preserve that independence by itself.
For a decoder-only Transformer, causal masking blocks attention to future positions. It does not normally block attention to earlier positions that belong to another packed example. If boundaries are ignored, tokens in a later example can attend to keys and values from an earlier one. The model then receives context that the data pipeline intended to keep separate.
Causality and example isolation are different constraints
Consider two tokenized examples packed into a single buffer:
A0 A1 A2 B0 B1 B2 B3A conventional causal mask permits position i to attend to positions j <= i. Under that rule, B0 can attend to A0, A1, and A2, because those positions are in its past.
That behavior is causally valid but violates example isolation. The desired mask needs both conditions:
allowed(i, j) =
(j <= i)
AND
(segment_id[i] == segment_id[j])With segment identifiers
0 0 0 1 1 1 1the resulting attention pattern is block-diagonal within the lower-triangular causal structure. Tokens from example B can attend to earlier tokens from B, but not to tokens from A.
The distinction matters because causal order is a property of token positions, while example membership is a property introduced by the batching and packing scheme. One constraint cannot stand in for the other.
Loss masking does not repair attention leakage
Training pipelines often mask selected labels so that padding, prompt tokens, or separators do not contribute directly to the loss. That mechanism controls which positions produce supervised loss terms. It does not control which keys and values a query can attend to.
Suppose the labels associated with a separator are assigned an ignore index. A later token can still attend through that separator to representations from the previous example unless the attention mechanism blocks the path. Removing a position from the loss therefore does not make its hidden state or preceding context invisible.
This separation is useful when reviewing a packing implementation. Label masks answer which predictions affect the objective. Attention masks answer which token states can influence each prediction. A correct packed batch can require both.
Position IDs are a separate design choice
Packing also raises a position-index question. A pipeline can keep positions increasing across the whole packed buffer:
0 1 2 3 4 5 6or reset them at each example boundary:
0 1 2 0 1 2 3Neither representation, by itself, enforces attention isolation. Position IDs affect positional information supplied to the model; they are not an access-control mask over keys and values.
The appropriate position scheme depends on the model architecture and the training setup. Absolute position embeddings, rotary position mechanisms, relative position methods, and custom attention kernels can interpret position metadata differently. A packing implementation therefore should not assume that resetting position IDs is equivalent to blocking cross-example attention.
The same caution applies in the opposite direction. A boundary-aware attention mask can isolate examples even when position IDs continue across the packed buffer, but that position convention may still differ from the distribution used elsewhere in training or inference.
Separator tokens do not create a hard boundary
A dedicated end-of-sequence or separator token can signal a boundary in the token stream. It does not create a structural attention barrier unless the model or attention mask gives it that behavior.
With only a causal mask, tokens after the separator remain able to attend to tokens before it. The model may choose to use the separator as a semantic cue, but that is different from making cross-boundary attention impossible.
This distinction becomes especially relevant when packed examples are unrelated. A separator can tell the model that one document ended, while the unmasked attention graph still exposes the next document to the previous document’s hidden states. If independence is part of the intended training example definition, a structural mask is the direct mechanism for enforcing it.
Packed masks do not have to be materialized as dense matrices
The conceptual mask is easy to express as an N x N boolean matrix, but materializing that matrix is not required by the semantics. Modern attention implementations may represent boundaries through sequence lengths, cumulative offsets, block metadata, or kernel-specific structures.
For example, a packed batch can describe contiguous sequence lengths:
lengths = [3, 4, 2]An attention kernel that explicitly supports variable-length packed sequences can use such metadata to keep attention within each range without storing a dense block-diagonal mask.
This is an implementation property, not a universal API guarantee. A tensor named attention_mask can mean a padding mask in one library, an additive score mask in another, or metadata consumed by a specialized kernel elsewhere. Correctness depends on the actual contract of the model and attention implementation.
The invariant is more stable than the representation: no query belonging to one independent example should receive key/value contributions from another independent example when isolation is required.
Boundary errors can survive superficial validation
A packed batch can have the expected tensor shape, token count, label count, and loss value while still using the wrong attention graph. Cross-example visibility does not necessarily trigger an exception or produce an obviously invalid scalar.
That makes structural validation more useful than shape checks alone. For a small synthetic packed batch, the attention permission pattern can be inspected against expected segment boundaries. If an implementation exposes only compact boundary metadata, tests can verify the offsets or sequence lengths supplied to the kernel.
The useful assertion is about connectivity. For every query-key pair from different independent segments, the effective attention path should be blocked. For pairs within one segment, ordinary causal constraints should still apply.
Packing changes representation, not sample semantics
Packing is often introduced as a batching optimization: fewer padded tokens occupy compute and memory. That optimization is valid only if the packed representation preserves the semantics required by the training objective.
Some workloads intentionally allow context to cross document boundaries. Others concatenate related records into a single training sequence on purpose. In those cases, blocking every boundary would encode a different objective. The boundary-aware rule applies when the packed units are intended to remain independent examples.
That condition is the implementation boundary to keep explicit. Packing several examples into one tensor does not automatically turn them into one context, and causal masking does not automatically keep them separate. The attention graph has to match the sample semantics chosen by the data pipeline.