Skip to content

Archive

DMA

8 articles
Tech 22 Sep 2026 6 min read

NVMe Completion Queue Phase Tags Distinguish New Entries After Ring Wrap

An NVMe completion queue is a circular memory structure shared by a controller and host software. The controller posts completion queue entries after commands finish, while the host consumes those entries and advances its queue head. Once either side reaches the final slot, its index wraps to slot zero. That wrap creates a small but important state problem. Queue memory still contains bytes from earlier completions. Reading a nonzero entry at the current head is not enough to prove that the controller has posted a fresh completion there. NVMe solves this with a one-bit Phase Tag carried in every completion queue entry.

Tech 19 Sep 2026 5 min read

Streaming DMA Mappings Transfer Buffer Ownership Between CPU and Device

Streaming DMA Mappings Transfer Buffer Ownership Between CPU and Device A DMA buffer can be valid memory for both a CPU and a device while still requiring a strict handoff between them. Linux streaming DMA mappings express that handoff. The mapping API supplies a device-visible DMA address and gives the DMA layer a point at which architecture-specific cache maintenance, address translation, or bounce buffering can occur. This matters most on systems where device DMA is not automatically coherent with CPU caches, but the ownership rules are part of the portable DMA API even on machines where cache maintenance becomes a no-op.

Tech 19 Sep 2026 5 min read

PCIe PASID Lets One Device Carry Multiple Address-Space Contexts

PCIe PASID Lets One Device Carry Multiple Address-Space Contexts A PCIe function normally has a Requester ID derived from its bus, device, and function identity. That identifier tells platform components which function issued a transaction, but it is too coarse when one device serves work from several process address spaces at the same time. Process Address Space ID (PASID) adds another identity field to the transaction path. A PASID can select an address-space context beneath the same device function, allowing the IOMMU to distinguish memory traffic that belongs to different processes without requiring a separate PCIe function for each one.

Tech 19 Sep 2026 5 min read

PCIe ATS Moves Address Translation Caching Into the Device

PCIe ATS Moves Address Translation Caching Into the Device An IOMMU can translate DMA addresses on behalf of a device, but that arrangement puts translation machinery in the path of device memory traffic. PCIe Address Translation Services (ATS) adds another option: a capable device can request a translation and retain the result in its own translation cache. Later transactions can carry the translated address instead of requiring the same translation work at the IOMMU for every access.

Tech 19 Sep 2026 6 min read

PCIe ACS Controls Whether Peer Traffic Can Bypass the IOMMU

PCIe ACS Controls Whether Peer Traffic Can Bypass the IOMMU An IOMMU can restrict DMA only for transactions that reach its translation and permission checks. PCI Express complicates that boundary because two endpoints under the same hierarchy can exchange peer-to-peer traffic without sending every transaction through the root complex. A switch may be able to route a request directly from one downstream port to another. PCIe Access Control Services (ACS) adds controls for that routing boundary. Depending on the component and supported ACS features, the fabric can validate a request, block selected translated traffic, redirect peer requests or completions upstream, or constrain their egress. The result is not merely a routing preference. ACS can determine whether two devices are separable at the IOMMU boundary.

Tech 19 Sep 2026 6 min read

IOMMU Translation Separates Device DMA Addresses from Physical Memory

A DMA-capable device can issue memory transactions without the CPU copying each payload. On systems with an IOMMU, the address carried by that device transaction does not have to be a host physical address. The IOMMU can translate a device-visible I/O virtual address into a physical page and reject accesses outside the configured mapping. That translation boundary changes both isolation and data movement. Drivers and operating systems can give a device a constrained address space, while the hardware must maintain translation state close enough to the I/O path to avoid turning every DMA request into a page-table walk.

Tech 19 Sep 2026 6 min read

IOMMU IOTLB Invalidation Controls When DMA Remapping Takes Effect

IOMMU IOTLB Invalidation Controls When DMA Remapping Takes Effect Changing an IOMMU page-table entry does not necessarily change the translation used by the next DMA request. An IOMMU can cache address translations in an I/O translation lookaside buffer, commonly called an IOTLB. Software must invalidate affected cached state when a mapping is removed or replaced, then observe the invalidation semantics required by that IOMMU before treating the old translation as retired.

Tech 17 Sep 2026 7 min read

IOMMU Remaps Device DMA Addresses for Memory Isolation

Direct memory access lets a device move data between itself and system memory without making the CPU copy every byte. Network adapters, storage controllers, GPUs, and other high-throughput devices rely on DMA to keep data moving efficiently. That capability also creates a protection problem. A device that can issue unrestricted memory transactions could read or overwrite physical pages belonging to the kernel, another process, or another virtual machine. An input-output memory management unit, commonly called an IOMMU, places address translation and access control between DMA-capable devices and physical memory.