Skip to content

Archive

GPU

2 articles
Artificial Intelligence 24 Sep 2026 5 min read

FlashAttention Tiles Exact Softmax Without Materializing the Score Matrix

Standard scaled dot-product attention forms a score matrix whose two long axes are sequence positions. For one attention head, S = QK^T / sqrt(d) P = softmax(S) O = PV the mathematical definition is compact, but a direct GPU implementation can write the large intermediate matrices S and P to high-bandwidth memory before reading them again. FlashAttention changes that dataflow. It processes blocks of queries, keys, and values in on-chip memory and carries enough row-wise softmax state to combine score tiles exactly.

Tech 11 Sep 2026 7 min read

What Hardware Acceleration Does in a Web Browser

A browser setting called hardware acceleration can sound as if it simply makes the internet faster. That isn’t quite what it does. Your connection can stay exactly the same while hardware acceleration changes how efficiently the computer draws pages, plays video, runs visual effects, or handles other graphics-heavy work. The basic idea is straightforward: instead of asking the main processor to do every kind of computation itself, the browser can hand suitable work to hardware designed for that job. On a typical computer, that often means using the graphics processing unit, or GPU.