Skip to content

Archive

PCIe

18 articles
Tech 22 Sep 2026 7 min read

NVMe Completion Queue Phase Tags Mark Reused Entries

An NVMe completion queue is a fixed-size circular array in host memory. The controller writes completion queue entries as commands finish, while host software consumes those entries and advances the queue head. Eventually both sides return to slots that already contain data from an earlier circuit of the ring. Reusing memory creates a small but important ambiguity. A slot can contain a perfectly formed completion entry even when the controller has not written a new completion there yet. Clearing every consumed entry would add memory traffic and still require careful coordination between the host and controller.

Tech 22 Sep 2026 6 min read

NVMe Completion Queue Phase Tags Distinguish New Entries After Ring Wrap

An NVMe completion queue is a circular memory structure shared by a controller and host software. The controller posts completion queue entries after commands finish, while the host consumes those entries and advances its queue head. Once either side reaches the final slot, its index wraps to slot zero. That wrap creates a small but important state problem. Queue memory still contains bytes from earlier completions. Reading a nonzero entry at the current head is not enough to prove that the controller has posted a fresh completion there. NVMe solves this with a one-bit Phase Tag carried in every completion queue entry.

Tech 19 Sep 2026 4 min read

Write-Combining Device Memory Can Merge CPU Stores Before I/O

Write-Combining Device Memory Can Merge CPU Stores Before I/O A sequence of CPU stores to device memory does not necessarily become an identical sequence of bus transactions. With a write-combining mapping, the processor may collect adjacent stores and emit a larger transfer later. That behavior suits framebuffer-like regions and other device buffers built for bulk writes, but it changes the ordering and transaction assumptions a driver can safely make. Linux exposes this mapping class through interfaces such as ioremap_wc(). It is distinct from the default ioremap() mapping used for ordinary control registers.

Tech 19 Sep 2026 7 min read

PCIe Posted Writes Separate CPU Completion from Device Visibility

PCIe Posted Writes Separate CPU Completion from Device Visibility An MMIO store can be complete from the CPU’s point of view while the corresponding write is still moving through the I/O path. PCI and PCIe memory writes are normally posted: the requester does not wait for a completion response for each write. Bridges and interconnect logic can accept the transaction and let the CPU continue before the endpoint has consumed it.

Tech 19 Sep 2026 6 min read

PCIe Posted MMIO Writes Can Outlive the CPU Store That Issued Them

A CPU can retire or complete an MMIO store before the corresponding PCIe Memory Write has reached the target device. The gap exists because PCIe Memory Write requests are posted: the requester sends them without waiting for a completion packet from the completer. That property is useful for throughput, but it creates an important boundary. A software-visible store instruction, an ordering barrier, and device observation of the write are not automatically the same event.

Tech 19 Sep 2026 5 min read

PCIe PASID Lets One Device Carry Multiple Address-Space Contexts

PCIe PASID Lets One Device Carry Multiple Address-Space Contexts A PCIe function normally has a Requester ID derived from its bus, device, and function identity. That identifier tells platform components which function issued a transaction, but it is too coarse when one device serves work from several process address spaces at the same time. Process Address Space ID (PASID) adds another identity field to the transaction path. A PASID can select an address-space context beneath the same device function, allowing the IOMMU to distinguish memory traffic that belongs to different processes without requiring a separate PCIe function for each one.

Tech 19 Sep 2026 5 min read

PCIe ATS Moves Address Translation Caching Into the Device

PCIe ATS Moves Address Translation Caching Into the Device An IOMMU can translate DMA addresses on behalf of a device, but that arrangement puts translation machinery in the path of device memory traffic. PCIe Address Translation Services (ATS) adds another option: a capable device can request a translation and retain the result in its own translation cache. Later transactions can carry the translated address instead of requiring the same translation work at the IOMMU for every access.

Tech 19 Sep 2026 6 min read

PCIe ASPM Link States Trade Idle Power for Exit Latency

A PCI Express link does not need to keep every transmitter and receiver block fully active when no packets are moving. Active State Power Management, or ASPM, lets a link enter lower-power states during idle periods and return to L0 when traffic resumes. The mechanism sits below application I/O. Software can issue the same storage, network, or device operation regardless of the current link state, but the first transaction after an idle interval can encounter extra delay as the link returns to active operation.

Tech 19 Sep 2026 6 min read

PCIe ACS Controls Whether Peer Traffic Can Bypass the IOMMU

PCIe ACS Controls Whether Peer Traffic Can Bypass the IOMMU An IOMMU can restrict DMA only for transactions that reach its translation and permission checks. PCI Express complicates that boundary because two endpoints under the same hierarchy can exchange peer-to-peer traffic without sending every transaction through the root complex. A switch may be able to route a request directly from one downstream port to another. PCIe Access Control Services (ACS) adds controls for that routing boundary. Depending on the component and supported ACS features, the fabric can validate a request, block selected translated traffic, redirect peer requests or completions upstream, or constrain their egress. The result is not merely a routing preference. ACS can determine whether two devices are separable at the IOMMU boundary.

Tech 19 Sep 2026 7 min read

MSI-X Per-Vector Masking Separates Interrupt Control Across Device Queues

MSI-X Per-Vector Masking Separates Interrupt Control Across Device Queues A PCI function using MSI-X can expose multiple interrupt vectors whose delivery state is controlled independently. Software can mask one MSI-X table entry while other enabled entries remain able to signal interrupts. That property matters for devices with multiple queues because interrupt control can follow the same partitioning as the I/O work instead of collapsing every notification source behind one device-wide interrupt state.

Tech 19 Sep 2026 7 min read

MSI-X Lets Device Queues Target Separate CPU Interrupt Paths

MSI-X Lets Device Queues Target Separate CPU Interrupt Paths A multiqueue PCIe device can move data through many queues at once, yet a single interrupt path would funnel completion handling back through one signal. MSI-X removes that device-wide bottleneck from the interrupt interface. Each allocated MSI-X entry represents an independently configurable message-signaled interrupt, so a driver can associate different queues or event classes with different Linux IRQs and CPU affinity policies.

Tech 17 Sep 2026 7 min read

PCIe ASPM Trades Link Wake Latency for Idle Power

A PCI Express link does not need to remain at full active power while no packets are moving. Active State Power Management, commonly called ASPM, lets compatible link partners enter lower-power link states during idle periods and return to active operation when traffic resumes. The tradeoff is direct: deeper idle states can save more energy, but leaving them takes time. A system therefore balances link power against the latency added to the next transfer.

Tech 17 Sep 2026 9 min read

NVMe Queue Pairs Separate Command Submission from Completion

NVMe Queue Pairs Separate Command Submission from Completion An NVMe solid-state drive does not need the CPU to hand each storage command directly to a device register and then wait for that command to finish. Instead, NVMe places command and completion records in queues held in host memory. The controller reads pending commands from submission queues and writes results to associated completion queues. That arrangement matches fast PCIe storage well. Modern SSD controllers can process many operations at once across flash channels, internal dies, and controller pipelines. A queue model lets software keep that parallel hardware busy while avoiding a long series of synchronous command handoffs.

Tech 16 Sep 2026 5 min read

PCIe Relaxed Ordering Lets Transactions Pass Within Ordering Rules

PCIe Relaxed Ordering Lets Transactions Pass Within Ordering Rules PCI Express carries requests and completions through switches, bridges, and endpoint logic that can have several transactions in flight at once. Strict ordering between every packet would make many independent transfers wait behind traffic that has no dependency on them. Relaxed Ordering provides a protocol signal that permits more reordering where the requester can tolerate it. The feature does not remove all ordering constraints. It marks a transaction as eligible for additional movement relative to other traffic, while PCIe ordering rules still define which combinations may pass. Software and device logic must only use the attribute when reordering cannot expose stale state or break a producer-consumer dependency.

Tech 16 Sep 2026 5 min read

PCIe Active State Power Management Trades Idle Power for Exit Latency

A PCI Express link does not need to stay at its fully active electrical state while no traffic is moving. Active State Power Management, commonly shortened to ASPM, lets compatible link partners place the link into lower-power states during idle periods. The practical tradeoff is simple: deeper idle states can save more power, but returning to active operation takes time. That exit delay becomes part of the latency seen when new traffic arrives.

Tech 16 Sep 2026 3 min read

NVMe Doorbell Registers Notify Controllers of Queue Progress

NVMe Doorbell Registers Notify Controllers of Queue Progress NVMe places submission and completion queues in host memory, but a controller still needs a signal when software adds commands or consumes completion entries. Doorbell registers provide that signal. Host software writes queue pointer values to memory-mapped controller registers so the device can track progress without scanning host memory continuously. The mechanism separates queue storage from queue notification. Commands and completion entries live in DMA-accessible memory, while small register writes tell the controller which portion of each queue has changed.

Tech 15 Sep 2026 6 min read

PCIe Link Width Sets How Many Lanes Carry Data

A PCI Express slot can look physically large while providing fewer active lanes than its connector suggests. A full-length slot may carry sixteen lanes, eight lanes, four lanes, or even one lane depending on the motherboard design and current resource allocation. PCIe calls these arrangements link widths. Common widths include x1, x4, x8, and x16. The number after the x indicates how many lanes participate in the link. That distinction matters for graphics cards, storage adapters, network cards, capture hardware, accelerators, and other devices that can move large amounts of data through an expansion slot.

Tech 15 Sep 2026 7 min read

NVMe Queues Let Storage Handle Many Commands in Parallel

NVMe storage does not send every read or write through one shared command line. The protocol is built around queue pairs: software places commands into a submission queue, and the controller reports finished work through a corresponding completion queue. That structure matters most when several processor cores and application threads are generating storage work at the same time. Multiple queues can distribute command handling across cores, reduce contention around a single software path, and keep a fast solid-state drive supplied with enough outstanding work.