FlashAttention Tiles Exact Softmax Without Materializing the Score Matrix
Standard scaled dot-product attention forms a score matrix whose two long axes are sequence positions. For one attention head, S = QK^T / sqrt(d) P = softmax(S) O = PV the mathematical definition is compact, but a direct GPU implementation can write the large intermediate matrices S and P to high-bandwidth memory before reading them again. FlashAttention changes that dataflow. It processes blocks of queries, keys, and values in on-chip memory and carries enough row-wise softmax state to combine score tiles exactly.