Padding Masks Do Not Remove Padding Compute in Dense Attention
A batch can contain two prompts with very different token counts yet represent both with the same rectangular tensor. The shorter prompt is extended with padding so its tensor shape matches the longest sequence in the batch. An attention mask can stop those padded positions from contributing to attention probabilities, but that semantic exclusion does not imply that dense kernels skip every operation associated with the padded rows and columns.