Skip to content

Archive

Memory Ordering

8 articles
Tech 23 Sep 2026 6 min read

PCIe Posted Writes Separate Transmission from Device Observation

PCI Express does not treat every request as a round trip. A Memory Write Request is normally a posted transaction: the requester sends a transaction layer packet carrying the write, but the completer does not return a Completion TLP for that individual request. This removes completion traffic from a common high-volume path, while placing more importance on ordering and explicit synchronization when software needs evidence that a write has progressed far enough.

Linux 23 Sep 2026 5 min read

membarrier Private Expedited Orders Userspace Memory Across Threads

Linux membarrier() can move part of a synchronization cost from a frequently executed path to a less frequent coordination path. With MEMBARRIER_CMD_PRIVATE_EXPEDITED, one thread asks the kernel to establish a memory-ordering point across the running threads in the same process. The caller pays for the system call when coordination is needed instead of requiring every fast-path execution to carry an explicit hardware memory barrier. This is narrower than a general thread rendezvous. membarrier() does not run an application callback on sibling threads, does not wait for application-level acknowledgements, and does not turn ordinary data races into valid synchronization. Its contract concerns ordering of userspace memory accesses around the barrier.

Tech 19 Sep 2026 4 min read

Write-Combining Device Memory Can Merge CPU Stores Before I/O

Write-Combining Device Memory Can Merge CPU Stores Before I/O A sequence of CPU stores to device memory does not necessarily become an identical sequence of bus transactions. With a write-combining mapping, the processor may collect adjacent stores and emit a larger transfer later. That behavior suits framebuffer-like regions and other device buffers built for bulk writes, but it changes the ordering and transaction assumptions a driver can safely make. Linux exposes this mapping class through interfaces such as ioremap_wc(). It is distinct from the default ioremap() mapping used for ordinary control registers.

Software Engineering 19 Sep 2026 7 min read

Sequence Counters Detect Concurrent Writes Without Reader Locks

A sequence counter can let readers copy shared state without taking the writer’s lock. The reader samples a counter, copies the protected fields, then samples the counter again. A stable even value at both observations indicates that no writer overlapped the copy under the synchronization contract. A changed or odd value forces the reader to discard the snapshot and retry. This pattern moves work away from reader-side lock ownership, but it does not remove synchronization. Writers still need serialization, counter transitions need defined memory-ordering semantics, and the protected data must remain safe to access during an overlapping write. Those constraints make sequence counters suitable for some read-mostly snapshots and unsafe for data whose lifetime can disappear beneath a reader.

Tech 19 Sep 2026 6 min read

PCIe Posted MMIO Writes Can Outlive the CPU Store That Issued Them

A CPU can retire or complete an MMIO store before the corresponding PCIe Memory Write has reached the target device. The gap exists because PCIe Memory Write requests are posted: the requester sends them without waiting for a completion packet from the completer. That property is useful for throughput, but it creates an important boundary. A software-visible store instruction, an ordering barrier, and device observation of the write are not automatically the same event.

Software Engineering 19 Sep 2026 6 min read

Linux membarrier Moves Memory Ordering Cost to a Coordinating Thread

A concurrent runtime can have thousands of fast-path operations for every rare state transition that requires global coordination. Placing a full memory barrier on every fast path makes each operation pay for that rare transition. Linux membarrier() supports the opposite arrangement: a coordinating thread enters the kernel and forces a defined ordering point across a target set of threads, moving more cost to the infrequent side of the protocol. This is not a generic replacement for atomics, mutexes, or language memory models. It is a Linux kernel interface whose guarantees apply to memory accesses and targeted threads under specific commands. Correct use requires a protocol that already defines which accesses occur before and after the coordination point.

Linux 18 Sep 2026 6 min read

membarrier Moves Memory-Ordering Cost to an Infrequent Coordination Path

A full hardware memory barrier in a frequently executed path can impose a cost on every operation, even when cross-thread coordination happens only occasionally. Linux membarrier() supports a different placement of that cost: a rare coordination path can request an ordering event across a defined set of threads while a frequent path may need only compiler-level ordering. This is not a general replacement for atomics, locks, or the memory model of a programming language. It is a Linux-specific synchronization primitive for designs whose correctness already has a precise pairing between a frequent path and an infrequent coordination path.

Tech 17 Sep 2026 6 min read

Memory Fences Constrain Cross-Core Memory Ordering

Memory Fences Constrain Cross-Core Memory Ordering A processor can execute memory operations with more freedom than source-code order suggests. Loads may begin early, stores may wait in buffers, cache-coherence traffic may complete at different times, and independent operations can overlap. These techniques improve throughput, but concurrent software needs precise rules for publishing and observing shared state. A memory fence places an ordering constraint around selected memory operations. It does not normally flush every cache, serialize the entire processor, or make all cores execute one instruction stream. Its role is narrower: it restricts which memory-order outcomes are permitted across a defined boundary.