Skip to content

Archive

FlashAttention

1 articles
Artificial Intelligence 24 Sep 2026 5 min read

FlashAttention Tiles Exact Softmax Without Materializing the Score Matrix

Standard scaled dot-product attention forms a score matrix whose two long axes are sequence positions. For one attention head, S = QK^T / sqrt(d) P = softmax(S) O = PV the mathematical definition is compact, but a direct GPU implementation can write the large intermediate matrices S and P to high-bandwidth memory before reading them again. FlashAttention changes that dataflow. It processes blocks of queries, keys, and values in on-chip memory and carries enough row-wise softmax state to combine score tiles exactly.