PCI Express does not treat every request as a round trip. A Memory Write Request is normally a posted transaction: the requester sends a transaction layer packet carrying the write, but the completer does not return a Completion TLP for that individual request. This removes completion traffic from a common high-volume path, while placing more importance on ordering and explicit synchronization when software needs evidence that a write has progressed far enough.

The distinction is especially relevant to device drivers. Writing a memory-mapped device register can cause a CPU store instruction to retire before the target device has acted on the new value. Between the core and the endpoint sit CPU ordering rules, host bridges, PCIe transaction queues, switches, and device-side buffering.

Posted means no transaction-layer completion for the write

PCIe classifies requests according to their completion behavior. Reads are non-posted because the requester needs returned data. Memory writes are posted because the request carries the payload and does not require an ordinary Completion TLP from the destination.

requester                         endpoint

Memory Write TLP  ------------------>
       |
       |       no Completion TLP
       |
subsequent work can continue

This property improves efficiency. A stream of writes does not consume reverse-path bandwidth with one completion packet per write, and the requester does not need to retain state waiting for those completions.

It also means the absence of a completion cannot be used as proof that the endpoint has consumed the write. Software requiring such a guarantee needs a protocol that creates an observable ordering point.

CPU completion and PCIe progress are different boundaries

A device register exposed through memory-mapped I/O looks like an address to software, but its behavior is not equivalent to ordinary cacheable RAM. A CPU can hand a store to the platform’s I/O path and continue execution while the transaction moves toward the endpoint.

Architecture-specific MMIO accessors and barriers exist partly to keep software from assuming stronger ordering than the platform provides. Driver code should use the operating system’s device-I/O interfaces rather than replacing them with arbitrary pointer stores and compiler assumptions.

The relevant boundary also depends on the operation being coordinated. A driver may only need two register writes to reach a device in a defined order, or it may need DMA descriptors in system memory to become visible before ringing a device doorbell. Those are related but distinct ordering problems involving CPU memory rules, DMA coherence, and PCIe transaction ordering.

A read can provide a synchronization point

A common pattern after posted MMIO writes is to perform a read from the same device when the driver needs to flush posted writes through the path. The read is non-posted and requires a Completion carrying the requested data. Under the applicable PCIe ordering rules, completing that read can establish that earlier writes ordered before it have progressed as required.

host                              device

write control  -------------------->
write doorbell -------------------->
read status    -------------------->
               <-------------------- Completion + data

continue after read completion

This is not a universal instruction to read any register. The chosen register must be safe to read, and the device specification can define side effects or special semantics for particular locations. Operating-system driver APIs may also provide dedicated helpers for flushing or ordered MMIO.

The important point is structural: a posted write supplies no per-request completion, while a subsequent non-posted operation can create a response dependency that software can use as part of an ordering protocol.

Ordering attributes can change permitted reordering

PCIe transaction ordering is governed by more than request type. Attributes such as Relaxed Ordering can permit traffic to pass other transactions in cases that default ordering would constrain. Traffic classes and virtual channels also affect how packets move through the fabric.

A driver cannot safely derive its synchronization model from packet direction alone. The device, platform, and operating-system interfaces define which accesses are ordered and which barriers are required around them.

For ordinary control paths, driver frameworks typically encode architecture-specific details behind primitives such as register read and write operations plus memory barriers. This keeps source code from depending on one CPU’s MMIO behavior when the same driver model must work across several architectures.

DMA publication adds another ordering layer

Doorbell registers often tell a device that new descriptors are ready in host memory. In that sequence, two visibility domains matter. First, descriptor stores must become visible to the device’s DMA reads. Second, the MMIO doorbell must not be observed in a manner that lets the device fetch descriptors before their contents are ready.

A simplified sequence is:

CPU fills descriptor in memory
        |
DMA write barrier
        |
MMIO write to device doorbell
        |
device fetches descriptor

The exact barrier depends on the operating system, architecture, mapping type, and DMA coherence model. A generic CPU memory barrier and an MMIO flush are not interchangeable concepts. One orders memory visibility; another may be used to force posted I/O progress through a bridge or fabric.

This separation matters during debugging. A device that occasionally sees stale descriptor fields can reflect missing DMA publication ordering even if the doorbell register itself eventually receives the expected value.

Error handling does not turn posted writes into acknowledged writes

PCIe includes error reporting mechanisms, link-level reliability, and transaction-layer rules, but those mechanisms do not create a normal Completion TLP for each posted Memory Write Request. Link-layer acknowledgment confirms successful transfer across a link segment; it is not an application-level statement that the endpoint has performed the semantic action associated with a register write.

Likewise, Advanced Error Reporting can surface classes of protocol and device errors without converting posted traffic into a request-response exchange. Driver protocols that require device confirmation usually obtain it through a status register, completion queue, interrupt, DMA-updated state, or another device-defined mechanism.

The distinction prevents a common category error: reliable packet delivery inside the PCIe fabric is not the same guarantee as completion of the operation represented by the packet payload.

Posted writes favor throughput but move synchronization into protocol design

Posted writes are efficient precisely because the common write path does not wait for a response packet. That property fits command submission, doorbells, configuration of runtime registers, and other write-heavy device interactions.

When software needs a stronger boundary, it has to use the ordering facilities supplied by the CPU architecture, operating system, PCIe rules, and device interface. Depending on the case, that can involve a memory barrier, an ordered MMIO accessor, a safe readback, or a device-level completion mechanism.

The useful mental model is to separate transmission from observation. Sending a posted Memory Write Request advances data toward the endpoint without an ordinary transaction completion. Establishing that later work may depend on the write requires a separate ordering or acknowledgment path matched to the device protocol.