Skip to content

Archive

Batching

2 articles
Artificial Intelligence 22 Sep 2026 6 min read

Padding Masks Do Not Remove Padding Compute in Dense Attention

A batch can contain two prompts with very different token counts yet represent both with the same rectangular tensor. The shorter prompt is extended with padding so its tensor shape matches the longest sequence in the batch. An attention mask can stop those padded positions from contributing to attention probabilities, but that semantic exclusion does not imply that dense kernels skip every operation associated with the padded rows and columns.

Artificial Intelligence 16 Sep 2026 6 min read

Bucket Sequence Lengths to Reduce Padding Waste

A padded batch is shaped by its longest sequence, not its average sequence. If one batch contains token counts of 120, 124, 131, and 900, every sequence may be represented at length 900. Most positions in the first three rows then carry padding rather than input tokens. Length bucketing changes batch composition instead of changing the model. Examples with similar token counts are placed near each other before batches are formed. The maximum length inside each batch falls closer to the lengths of its members, reducing the number of padded positions processed by operations that still use the rectangular batch shape.