Skip to content

Archive

I/O

37 articles
Tech 23 Sep 2026 6 min read

PCIe Posted Writes Separate Transmission from Device Observation

PCI Express does not treat every request as a round trip. A Memory Write Request is normally a posted transaction: the requester sends a transaction layer packet carrying the write, but the completer does not return a Completion TLP for that individual request. This removes completion traffic from a common high-volume path, while placing more importance on ordering and explicit synchronization when software needs evidence that a write has progressed far enough.

Linux 23 Sep 2026 4 min read

EPOLLEXCLUSIVE Limits Wakeups Across epoll Instances

EPOLLEXCLUSIVE Limits Wakeups Across epoll Instances EPOLLEXCLUSIVE changes event distribution when several epoll instances register the same target file description. Without the flag, a readiness event can wake waiters associated with every interested epoll instance. With exclusive registrations, Linux can restrict that fan-out and reduce redundant wakeups. The flag does not make delivery single-consumer, does not assign ownership of the target, and does not replace application-level coordination. Its contract is narrower: among epoll instances that registered a target with EPOLLEXCLUSIVE, one or more receive an event for a wakeup rather than requiring all of them to receive it.

Tech 23 Sep 2026 7 min read

epoll Readiness Modes Change How Event Loops Drain File Descriptors

Linux epoll lets one thread wait on readiness changes across many file descriptors without scanning every descriptor on each iteration. The interface is common in network servers, proxies, runtimes, and other programs that keep large sets of sockets active. The registration mode matters. Level-triggered operation keeps reporting a descriptor while the relevant condition remains ready. Edge-triggered operation reports transitions in readiness and expects the application to consume available work until the descriptor would block.

Tech 22 Sep 2026 7 min read

NVMe Completion Queue Phase Tags Mark Reused Entries

An NVMe completion queue is a fixed-size circular array in host memory. The controller writes completion queue entries as commands finish, while host software consumes those entries and advances the queue head. Eventually both sides return to slots that already contain data from an earlier circuit of the ring. Reusing memory creates a small but important ambiguity. A slot can contain a perfectly formed completion entry even when the controller has not written a new completion there yet. Clearing every consumed entry would add memory traffic and still require careful coordination between the host and controller.

Tech 22 Sep 2026 4 min read

io_uring Registered Buffers Pin User Pages for Repeated I/O

Linux io_uring can submit asynchronous I/O with ordinary user buffers, but repeated operations may still require the kernel to resolve and pin the relevant user pages for each request. Registered buffers move part of that work into an explicit setup phase. An application registers one or more memory regions with the ring. The kernel records those regions and keeps the backing pages pinned while the registration remains active. Later requests can refer to a registered region by index instead of presenting an arbitrary buffer that must be prepared from scratch.

Tech 19 Sep 2026 5 min read

Streaming DMA Mappings Transfer Buffer Ownership Between CPU and Device

Streaming DMA Mappings Transfer Buffer Ownership Between CPU and Device A DMA buffer can be valid memory for both a CPU and a device while still requiring a strict handoff between them. Linux streaming DMA mappings express that handoff. The mapping API supplies a device-visible DMA address and gives the DMA layer a point at which architecture-specific cache maintenance, address translation, or bounce buffering can occur. This matters most on systems where device DMA is not automatically coherent with CPU caches, but the ownership rules are part of the portable DMA API even on machines where cache maintenance becomes a no-op.

Tech 19 Sep 2026 7 min read

PCIe Posted Writes Separate CPU Completion from Device Visibility

PCIe Posted Writes Separate CPU Completion from Device Visibility An MMIO store can be complete from the CPU’s point of view while the corresponding write is still moving through the I/O path. PCI and PCIe memory writes are normally posted: the requester does not wait for a completion response for each write. Bridges and interconnect logic can accept the transaction and let the CPU continue before the endpoint has consumed it.

Tech 19 Sep 2026 6 min read

PCIe ASPM Link States Trade Idle Power for Exit Latency

A PCI Express link does not need to keep every transmitter and receiver block fully active when no packets are moving. Active State Power Management, or ASPM, lets a link enter lower-power states during idle periods and return to L0 when traffic resumes. The mechanism sits below application I/O. Software can issue the same storage, network, or device operation regardless of the current link state, but the first transaction after an idle interval can encounter extra delay as the link returns to active operation.

Tech 19 Sep 2026 7 min read

MSI-X Per-Vector Masking Separates Interrupt Control Across Device Queues

MSI-X Per-Vector Masking Separates Interrupt Control Across Device Queues A PCI function using MSI-X can expose multiple interrupt vectors whose delivery state is controlled independently. Software can mask one MSI-X table entry while other enabled entries remain able to signal interrupts. That property matters for devices with multiple queues because interrupt control can follow the same partitioning as the I/O work instead of collapsing every notification source behind one device-wide interrupt state.

Tech 19 Sep 2026 7 min read

MSI-X Lets Device Queues Target Separate CPU Interrupt Paths

MSI-X Lets Device Queues Target Separate CPU Interrupt Paths A multiqueue PCIe device can move data through many queues at once, yet a single interrupt path would funnel completion handling back through one signal. MSI-X removes that device-wide bottleneck from the interrupt interface. Each allocated MSI-X entry represents an independently configurable message-signaled interrupt, so a driver can associate different queues or event classes with different Linux IRQs and CPU affinity policies.

Tech 19 Sep 2026 6 min read

IOMMU IOTLB Invalidation Controls When DMA Remapping Takes Effect

IOMMU IOTLB Invalidation Controls When DMA Remapping Takes Effect Changing an IOMMU page-table entry does not necessarily change the translation used by the next DMA request. An IOMMU can cache address translations in an I/O translation lookaside buffer, commonly called an IOTLB. Software must invalidate affected cached state when a mapping is removed or replaced, then observe the invalidation semantics required by that IOMMU before treating the old translation as retired.

Software Engineering 19 Sep 2026 5 min read

EPOLLET Makes Readiness a State-Transition Contract

A descriptor registered with EPOLLET can remain readable after an event has been delivered without appearing again in the next epoll_wait(). The kernel reports a readiness transition; it does not promise to repeat the same notification merely because unread data remains. That distinction turns edge-triggered epoll into a state-transition contract between the kernel and the event loop. The consequence is structural. A handler cannot treat one event as permission for one read() and then return to the wait loop. With edge-triggered monitoring, the handler must account for all immediately available I/O state before relying on another transition.

Software Engineering 19 Sep 2026 5 min read

Edge-Triggered epoll Requires Draining Readiness to EAGAIN

With EPOLLET, an event loop can consume one notification, read only part of the available data, and then wait indefinitely even though unread bytes remain in the socket buffer. The descriptor is still ready, but no new readiness transition has occurred to generate another edge. That behavior makes edge-triggered epoll a contract between notification semantics and nonblocking I/O. The event says that readiness changed; it is not a promise that the kernel will keep repeating the same notification until the application finishes the work.

Software Engineering 18 Sep 2026 7 min read

Linux splice Makes Pipe Capacity Part of Data-Transfer Semantics

Linux splice() can transfer bytes between file descriptors without routing those bytes through a user-space buffer, but the interface is not a generic descriptor-to-descriptor copy primitive. At least one endpoint must be a pipe. That requirement makes pipe state part of the transfer contract: capacity, readable data, writer presence, blocking mode, and partial progress can all affect an otherwise straightforward data path. The useful boundary is therefore not simply “kernel copy versus user copy.” splice() changes the shape of ownership and flow control. Application code stops owning an intermediate byte array, while it still owns the control loop that accounts for bytes transferred, handles readiness, and preserves offset semantics.

Software Engineering 18 Sep 2026 8 min read

Linux copy_file_range Separates Copy Semantics From Data Movement

copy_file_range() asks Linux to copy bytes between regular files without requiring the application to shuttle those bytes through a user-space buffer. The call defines a byte-range operation, but it does not prescribe the physical transfer mechanism. A filesystem can perform ordinary data movement, use a copy-on-write sharing mechanism such as reflink, or employ another supported acceleration path while preserving the visible file contents required by the operation. That separation is the central API boundary. Applications specify source and destination ranges and observe the returned byte count. The kernel and filesystem retain latitude over the mechanism used to realize the copy.

Software Engineering 18 Sep 2026 6 min read

EPOLLEXCLUSIVE Limits Wakeups Across Competing epoll Instances

EPOLLEXCLUSIVE changes which epoll waiters are awakened when several epoll instances monitor the same target. Without the flag, a readiness event can be delivered to every attached epoll instance. With exclusive registration, Linux can wake a smaller subset, reducing redundant scheduling in configurations that otherwise create a thundering herd. The flag changes wakeup distribution. It does not assign permanent ownership of the target descriptor, serialize I/O, or guarantee that exactly one application thread consumes each unit of work.

Linux 18 Sep 2026 4 min read

EPOLLEXCLUSIVE Limits Wakeups Across Competing epoll Instances

A single ready socket can wake several threads when each thread waits on a different epoll instance that watches that socket. Linux provides EPOLLEXCLUSIVE to narrow that wakeup fan-out: among epoll instances that registered the target with the flag, a readiness event wakes one or more rather than all of them. The distinction is deliberately weaker than “exactly one waiter.” EPOLLEXCLUSIVE changes notification selection across epoll instances. It does not transfer ownership of the file descriptor, serialize all I/O, or guarantee that only one thread can observe useful work.

Tech 17 Sep 2026 6 min read

Write Combining Merges Adjacent Stores Before Memory Traffic

Some memory regions are written far more often than they are read. Frame buffers, device apertures, and streaming output areas are common examples. Sending every small CPU store as a separate memory transaction can waste bus bandwidth and transaction overhead. Write combining gives the processor a temporary place to collect compatible stores. Several writes targeting nearby addresses can be merged into a larger transaction before they leave the CPU. The technique favors sustained write throughput, but it changes the timing and ordering properties that software can safely assume.

Software Engineering 17 Sep 2026 7 min read

Linux splice Makes the Pipe a Kernel Data-Transfer Boundary

splice() can transfer bytes between two file descriptors without first copying the payload into a userspace buffer, but the interface requires at least one endpoint to be a pipe. That requirement makes the pipe more than an incidental transport. It is the kernel-visible buffer boundary around which the operation’s offset, blocking, capacity, and partial-progress semantics are defined. This differs from a conventional read() followed by write(). In that sequence, userspace owns an intermediate byte array and can inspect or modify it. With splice(), the payload can remain in kernel-managed storage while the process coordinates movement between endpoints.

Software Engineering 17 Sep 2026 5 min read

Linux io_uring Separates Submission From Completion Ownership

Linux io_uring Separates Submission From Completion Ownership An io_uring request can remain in flight after the application has finished constructing its submission queue entry. That creates a lifetime boundary absent from a simple synchronous call: request metadata may become stable at submission, while memory used as the actual I/O payload can still be accessed until the operation completes. The completion queue is therefore not only a result channel. For many operations, it marks the point at which application-owned operation state can be reclaimed or reused.

Software Engineering 17 Sep 2026 6 min read

Linux epoll Edge Triggering Reports Readiness Transitions, Not Work Units

Linux epoll Edge Triggering Reports Readiness Transitions, Not Work Units With EPOLLET, an epoll interest does not behave like a queue containing one event for each byte, packet, connection, or application message. It reports changes in readiness state. Once a file descriptor is ready, additional work can accumulate without producing another edge that an application may rely on. The handler therefore has to consume available work until the nonblocking operation reports that progress would block.

Software Engineering 17 Sep 2026 8 min read

Linux copy_file_range Separates Copy Semantics From Copy Implementation

Linux copy_file_range Separates Copy Semantics From Copy Implementation A successful copy_file_range() call reports a byte count, not a promise about the physical path those bytes took. Linux can satisfy the request through filesystem-specific acceleration, an in-kernel transfer path, or another implementation permitted by the active filesystem interfaces. The application receives a range-copy operation with defined offset and return-value semantics; it does not receive a guarantee that storage blocks were physically duplicated.

Software Engineering 17 Sep 2026 8 min read

io_uring Shared Rings Make Memory Ordering Part of the ABI

An io_uring queue is shared memory with two independent execution domains changing its state. User space prepares submission entries and advances queue metadata; the kernel consumes those submissions and later publishes completion entries. The ring layout removes a copy boundary, but it also makes memory visibility part of the interface contract. A plain source-level assignment to a queue tail is not sufficient as a portable model of publication. The entry data must become visible before the tail value that makes the entry eligible for consumption. On the completion side, user space must observe the kernel’s publication of a completion before reading fields from that completion. The ordering relation is part of correctness, not merely an optimization detail.

Linux 17 Sep 2026 4 min read

Duplicated File Descriptors Share an Open File Description

Two file descriptor numbers can move the same file offset. On Linux, this occurs when both descriptors refer to one open file description, as happens after dup() and across inherited descriptors after fork(). The distinction matters because a file descriptor is a process-visible integer, while the open file description is the kernel object that carries state for an open instance of a file. Treating those layers as interchangeable can produce offset interference, status-flag changes that cross descriptor boundaries, and surprising behavior after process creation.