Skip to content

Archive

CPU

13 articles
Tech 22 Sep 2026 7 min read

NUMA First-Touch Placement Ties Physical Pages to the Faulting CPU

On a NUMA machine, reserving virtual address space does not necessarily choose the physical NUMA node that will supply every page. For anonymous memory, physical allocation commonly happens later, when a CPU first faults on a page. Under the default local allocation policy, that fault can make initialization order part of the memory-placement decision. This behavior is often called first-touch placement. The thread that first writes or otherwise faults a page can cause Linux to allocate backing memory near the NUMA node on which that thread is running, subject to the active memory policy, cpuset constraints, available memory, and kernel fallback behavior.

Tech 19 Sep 2026 4 min read

Write-Combining Device Memory Can Merge CPU Stores Before I/O

Write-Combining Device Memory Can Merge CPU Stores Before I/O A sequence of CPU stores to device memory does not necessarily become an identical sequence of bus transactions. With a write-combining mapping, the processor may collect adjacent stores and emit a larger transfer later. That behavior suits framebuffer-like regions and other device buffers built for bulk writes, but it changes the ordering and transaction assumptions a driver can safely make. Linux exposes this mapping class through interfaces such as ioremap_wc(). It is distinct from the default ioremap() mapping used for ordinary control registers.

Tech 19 Sep 2026 6 min read

TLB Shootdowns Extend Page-Table Changes Across CPUs

TLB Shootdowns Extend Page-Table Changes Across CPUs Changing a page-table entry in memory does not by itself retire every translation derived from that entry. A CPU that previously used the mapping can retain it in a translation lookaside buffer, or TLB. On a multiprocessor system, other CPUs may hold their own cached copies, so a mapping change can require coordination beyond the CPU that modified the page table. Linux exposes this distinction through its TLB-flush interfaces. After page-table state changes, architecture code must make the affected translations unusable on every relevant CPU before software relies on the new mapping or releases memory that the old mapping could reach.

Tech 19 Sep 2026 7 min read

TLB Shootdowns Coordinate Page-Table Changes Across CPUs

A page-table entry can change in memory while another CPU still holds the old address translation in its translation lookaside buffer (TLB). Updating the page table alone therefore does not necessarily make the new mapping effective on every processor that has executed the affected address space. Operating systems close that gap with TLB invalidation. When a mapping change can make a cached translation unsafe, processors that may retain the translation must invalidate it before the kernel treats the change as globally complete. On a multiprocessor system, coordinating those remote invalidations is commonly called a TLB shootdown.

Tech 19 Sep 2026 7 min read

MSI-X Lets Device Queues Target Separate CPU Interrupt Paths

MSI-X Lets Device Queues Target Separate CPU Interrupt Paths A multiqueue PCIe device can move data through many queues at once, yet a single interrupt path would funnel completion handling back through one signal. MSI-X removes that device-wide bottleneck from the interrupt interface. Each allocated MSI-X entry represents an independently configurable message-signaled interrupt, so a driver can associate different queues or event classes with different Linux IRQs and CPU affinity policies.

Tech 19 Sep 2026 5 min read

Cache-Line False Sharing Moves Coherence Ownership Between CPUs

Cache-Line False Sharing Moves Coherence Ownership Between CPUs Two threads can update different variables without sharing a lock or touching the same bytes and still interfere at the hardware level. If those variables occupy the same cache line, a coherent multiprocessor treats their storage as one coherence unit. Repeated writes from different CPUs can therefore move ownership of that line between caches even though the program considers the variables independent.

Tech 17 Sep 2026 6 min read

Write Combining Merges Adjacent Stores Before Memory Traffic

Some memory regions are written far more often than they are read. Frame buffers, device apertures, and streaming output areas are common examples. Sending every small CPU store as a separate memory transaction can waste bus bandwidth and transaction overhead. Write combining gives the processor a temporary place to collect compatible stores. Several writes targeting nearby addresses can be merged into a larger transaction before they leave the CPU. The technique favors sustained write throughput, but it changes the timing and ordering properties that software can safely assume.

Tech 17 Sep 2026 6 min read

NUMA Makes Memory Location Part of Access Cost

NUMA Makes Memory Location Part of Access Cost A large multiprocessor server can expose one physical address space while giving different processors different paths to that memory. A load from a page attached to the processor running a thread can take a shorter route than a load from memory attached to another processor package or NUMA node. This arrangement is called non-uniform memory access, or NUMA. It lets systems scale memory capacity and bandwidth across multiple processor sockets or chiplet groups without forcing every memory request through one centralized controller.

Tech 17 Sep 2026 6 min read

Memory Fences Constrain Cross-Core Memory Ordering

Memory Fences Constrain Cross-Core Memory Ordering A processor can execute memory operations with more freedom than source-code order suggests. Loads may begin early, stores may wait in buffers, cache-coherence traffic may complete at different times, and independent operations can overlap. These techniques improve throughput, but concurrent software needs precise rules for publishing and observing shared state. A memory fence places an ordering constraint around selected memory operations. It does not normally flush every cache, serialize the entire processor, or make all cores execute one instruction stream. Its role is narrower: it restricts which memory-order outcomes are permitted across a defined boundary.

Tech 16 Sep 2026 7 min read

TLB Caches Recent Virtual Address Translations

Modern processors commonly execute programs in virtual address spaces. A load or store can begin with a virtual address while the memory system ultimately needs a physical location and access permissions. Page tables hold the mapping information, but consulting their hierarchy for every memory reference would add substantial work. A translation lookaside buffer, or TLB, keeps recently used address translations near the processor. A TLB hit supplies cached mapping information without a full page-table walk. A TLB miss triggers additional translation work even when the requested application data is already present in a CPU cache.

Tech 16 Sep 2026 7 min read

CPU Thermal Throttling Reduces Clock Speed at Temperature Limits

A processor can run at a high clock rate only while electrical, power, and temperature limits permit it. Heavy work raises transistor switching activity, which raises power consumption and heat output. If cooling cannot remove that heat quickly enough, the processor approaches a thermal limit. Modern CPUs respond automatically. Control logic reduces performance states so the chip generates less heat. This behavior is commonly called thermal throttling. Thermal throttling is not the same as a processor simply running below its advertised maximum frequency. Boost algorithms already vary clock speed according to workload, active core count, current, power, and temperature. Thermal throttling refers specifically to temperature becoming a limiting condition.

Tech 16 Sep 2026 6 min read

CPU Store Buffers Decouple Retirement from Cache Writes

A processor does not need every store instruction to finish its cache update before later instructions make progress. Modern cores commonly place completed stores into a store buffer, allowing the instruction to retire while the memory subsystem handles the write afterward. This separation improves throughput because cache ownership, coherence traffic, and other memory activity can take longer than the execution pipeline can afford to wait. The buffer acts as a queue between architectural execution and the cache hierarchy.

Tech 16 Sep 2026 6 min read

CPU Cache Associativity Limits Where Lines Can Reside

Processor caches keep recently used memory close to execution cores, but cache capacity alone does not determine which data can remain resident. Most general-purpose CPU caches divide storage into sets and give each set a fixed number of slots, commonly called ways. A memory block maps to a particular set. It can occupy any way inside that set, but it cannot move into an unrelated set merely because that other set has free space. This placement rule makes hardware lookup practical and fast, while creating a distinct source of misses when too many active blocks compete for the same set.