Skip to content

Archive

Linux

231 articles
Tech 22 Sep 2026 4 min read

io_uring Registered Buffers Pin User Pages for Repeated I/O

Linux io_uring can submit asynchronous I/O with ordinary user buffers, but repeated operations may still require the kernel to resolve and pin the relevant user pages for each request. Registered buffers move part of that work into an explicit setup phase. An application registers one or more memory regions with the ring. The kernel records those regions and keeps the backing pages pinned while the registration remains active. Later requests can refer to a registered region by index instead of presenting an arbitrary buffer that must be prepared from scratch.

Linux 22 Sep 2026 5 min read

EFD_SEMAPHORE Makes eventfd Reads Consume One Counter Unit

An eventfd normally turns its entire nonzero counter into one read result and resets the counter to zero. Creating it with EFD_SEMAPHORE changes only the read side: each successful read returns the 64-bit value 1 and subtracts one from the kernel-maintained counter. That difference lets several units accumulated by writers remain separately consumable. The object is still an eventfd, with the same counter, write rules, descriptor lifetime, and readiness integration.

Linux 21 Sep 2026 5 min read

openat2 Resolve Flags Constrain Path Traversal During Lookup

A pathname passed to openat() is resolved by the kernel, but the caller has limited control over traversal through intermediate components. Linux openat2() adds a resolve field that applies constraints to the complete lookup operation. The restriction is evaluated while components are traversed, rather than by validating a pathname in userspace and opening it later. That distinction matters when path components can change concurrently. A userspace sequence that checks a path and then opens it creates separate observations of mutable filesystem state. openat2() places the selected lookup policy in the same system call that returns the file descriptor.

Linux 21 Sep 2026 6 min read

memfd File Seals Freeze Shared Memory State

A memfd_create() file starts as a mutable anonymous file. It can be resized, written, and mapped much like a regular file, while its storage remains volatile and disappears after the last reference is released. File sealing adds a different phase to that lifecycle: after data has been populated, the kernel can permanently reject selected classes of later modification. That transition is useful when one process prepares bytes and then hands the same file description to another process. The receiver can inspect the seals attached to the inode instead of relying only on a convention that the sender will stop changing the object.

Linux 21 Sep 2026 6 min read

io_uring Multishot Accept Keeps One Accept Request Active

A normal accept request has a one-to-one shape: one submitted operation eventually yields one completion. io_uring multishot accept changes that relationship. A single accept SQE can remain active across multiple incoming connections and emit a separate completion queue entry for each accepted socket. The request is persistent, but not permanent. Each CQE carries enough state for the application to tell whether the original request can produce another completion. That boundary matters because a server that treats every successful CQE as proof that accept is still armed can silently stop accepting after the multishot request terminates.

Linux 21 Sep 2026 4 min read

futex_waitv Blocks on Multiple Futex Words in One Wait

A traditional futex wait names one futex word. That maps cleanly to a mutex or condition whose blocking state is represented by one shared 32-bit value. Some synchronization designs instead need a thread to sleep until any member of several independent states changes. futex_waitv() provides that vector wait. The caller supplies an array of wait descriptors, each containing a futex address and an expected value. The kernel checks the vector and blocks only while every entry still matches its expected state.

Linux 20 Sep 2026 5 min read

pidfd_getfd Duplicates a Live File Descriptor Across Processes

A process can acquire a usable duplicate of a file descriptor that is already open in another process without asking that process to send it over a UNIX domain socket. Linux pidfd_getfd() performs that transfer through a PID file descriptor, subject to a ptrace access check. The returned descriptor is new in the caller, but the kernel object behind it is not independent. It refers to the same open file description as the target descriptor. That distinction controls offset sharing, file status flags, and operations on the underlying object.

Tech 19 Sep 2026 4 min read

Write-Combining Device Memory Can Merge CPU Stores Before I/O

Write-Combining Device Memory Can Merge CPU Stores Before I/O A sequence of CPU stores to device memory does not necessarily become an identical sequence of bus transactions. With a write-combining mapping, the processor may collect adjacent stores and emit a larger transfer later. That behavior suits framebuffer-like regions and other device buffers built for bulk writes, but it changes the ordering and transaction assumptions a driver can safely make. Linux exposes this mapping class through interfaces such as ioremap_wc(). It is distinct from the default ioremap() mapping used for ordinary control registers.

Linux 19 Sep 2026 5 min read

userfaultfd Write Protection Turns Memory Writes into Userspace Events

A thread can reach a valid, resident page and still stop before modifying it. With Linux userfaultfd write-protect mode, a registered page can be marked so that a write generates a userfaultfd page-fault event. A userspace handler receives that event, performs its bookkeeping, removes the protection, and lets the blocked thread continue. The mechanism sits between ordinary page permissions and application-level memory accounting. The page remains part of the process address space; the kernel redirects the write fault into a file-descriptor protocol instead of forcing the application to build the same control path around mprotect() and SIGSEGV.

Cybersecurity 19 Sep 2026 6 min read

userfaultfd Write Protection Moves Memory Writes Behind a User-Space Fault Boundary

A thread can hold a writable virtual memory mapping and still block when it attempts to modify a particular page. Linux userfaultfd write-protect mode lets user space register a memory range, apply write protection to pages in that range, and receive a page-fault event when a protected page is written. The VMA can remain logically writable while page-table state creates a narrower interception boundary. This mechanism is not a general authorization system. It is a memory-fault control interface. Its security relevance comes from the placement of the decision point: a write can be suspended before the protected page changes, allowing a separate handler to record state, coordinate migration, preserve a snapshot boundary, or reject progress by leaving the fault unresolved.

Linux 19 Sep 2026 5 min read

userfaultfd Moves Selected Page-Fault Resolution into User Space

A memory access normally enters the kernel page-fault path and completes without an application choosing the page contents at that instant. Linux userfaultfd changes that boundary for registered virtual address ranges: selected faults become events on a file descriptor, and a user-space manager can supply or activate the page before the faulting thread continues. The mechanism does not replace the process page tables with a user-space data structure. The kernel still owns page-table state and performs the final mapping operation. User space gains control over specific fault classes and the timing of their resolution.

Tech 19 Sep 2026 6 min read

TLB Shootdowns Extend Page-Table Changes Across CPUs

TLB Shootdowns Extend Page-Table Changes Across CPUs Changing a page-table entry in memory does not by itself retire every translation derived from that entry. A CPU that previously used the mapping can retain it in a translation lookaside buffer, or TLB. On a multiprocessor system, other CPUs may hold their own cached copies, so a mapping change can require coordination beyond the CPU that modified the page table. Linux exposes this distinction through its TLB-flush interfaces. After page-table state changes, architecture code must make the affected translations unusable on every relevant CPU before software relies on the new mapping or releases memory that the old mapping could reach.

Linux 19 Sep 2026 5 min read

timerfd Reads Count Periodic Expirations

A periodic timer can expire several times before a busy event loop gets CPU time again. Linux timerfd does not compress that delay into a bare “timer fired” notification. A successful read() returns an unsigned 64-bit count of expirations accumulated since the timer was armed or since the preceding successful read. That counter changes the semantics of delayed timer handling. Readiness says at least one expiration is pending; the value read from the descriptor says how many periods elapsed.

Linux 19 Sep 2026 5 min read

systemd Unit File Search Paths and Override Precedence

A systemd service file does not have to live in /etc/systemd/system. The system manager searches several directories for unit files, and the directory matters because the search order defines which definition wins when the same unit name exists in more than one place. For an administrator-managed service, /etc/systemd/system is usually the appropriate location. Distribution packages normally install units under /usr/lib/systemd/system, while /run/systemd/system is used for runtime configuration that disappears after reboot. This separation lets local configuration override vendor defaults without editing files owned by the package manager.

Tech 19 Sep 2026 5 min read

Streaming DMA Mappings Transfer Buffer Ownership Between CPU and Device

Streaming DMA Mappings Transfer Buffer Ownership Between CPU and Device A DMA buffer can be valid memory for both a CPU and a device while still requiring a strict handoff between them. Linux streaming DMA mappings express that handoff. The mapping API supplies a device-visible DMA address and gives the DMA layer a point at which architecture-specific cache maintenance, address translation, or bounce buffering can occur. This matters most on systems where device DMA is not automatically coherent with CPU caches, but the ownership rules are part of the portable DMA API even on machines where cache maintenance becomes a no-op.

Linux 19 Sep 2026 4 min read

signalfd Converts Linux Signals into File-Descriptor Events

A conventional POSIX signal can interrupt a thread at almost any instruction boundary and transfer control to a signal handler. That asynchronous control flow imposes strict limits on handler code and complicates programs whose main control plane already runs through epoll, poll, or select. Linux signalfd() provides a different delivery interface. A process blocks selected signals in the normal signal mask, creates a signalfd for that set, and receives pending signals by reading structured signalfd_siginfo records from the descriptor. The signal remains a signal at the kernel interface; only its consumption moves into ordinary file-descriptor I/O.

Linux 19 Sep 2026 5 min read

seccomp User Notifications Move Selected System Calls to a Supervisor

seccomp User Notifications Move Selected System Calls to a Supervisor A seccomp filter can do more than allow a system call or reject it in the kernel. When a filter returns SECCOMP_RET_USER_NOTIF, Linux can suspend the calling thread and deliver a description of that attempted system call to a user-space supervisor. The mechanism creates an interposition boundary around selected calls. It is useful when a less-privileged process needs an operation mediated by another process, such as a container manager handling a call that the container cannot perform directly. The boundary is deliberately narrower than a general security-policy engine: notification state can race with mutable target memory, and the kernel documentation warns against treating the supervisor’s inspection as an authorization primitive.

Software Engineering 19 Sep 2026 6 min read

seccomp User Notification Moves Selected Syscall Decisions to a Broker

A seccomp filter can do more than allow or reject a system call immediately. With user notification, a matching call can be suspended while another process receives a structured request on a listener file descriptor and decides what result the blocked thread receives. The mechanism turns selected syscall decisions into a brokered interface without moving the entire syscall implementation into user space. The boundary is precise but narrower than a general interposition layer. The kernel still owns syscall dispatch, task state, descriptor tables, and validation performed by kernel code. The broker receives metadata and can return a value, an error, or in supported cases request continued execution of the original syscall. Correct designs account for mutable target memory, notification lifetime, and the fact that a policy decision is not automatically a transaction over process state.

Linux 19 Sep 2026 4 min read

Random Hex Is Not a Hash: Generating Cryptographic Random Values on Linux

A command such as openssl rand -hex 32 is often described as generating a “random hash.” The output certainly looks like a SHA-256 digest: 64 hexadecimal characters. But no hash operation has happened. OpenSSL generated 32 random bytes and encoded them as hexadecimal. That distinction matters. A hash function transforms input into a fixed-size digest. A cryptographically secure random number generator produces unpredictable bytes. If the requirement is a token, session secret, API credential, nonce, or other fresh random value, the random bytes are the important part. Hashing them afterward usually adds no useful unpredictability.

Software Engineering 19 Sep 2026 5 min read

process_vm_readv and process_vm_writev Transfer Memory Across Process Boundaries

process_vm_readv() and process_vm_writev() let one Linux process copy bytes directly between its address space and another process’s address space. The calls operate on vectors of local and remote memory ranges, but a successful process lookup does not make remote memory stable. Mapping changes, page accessibility, permissions, and concurrent mutation remain separate parts of the contract. These interfaces are Linux-specific system calls. They do not define C object lifetime, synchronization, or a portable interprocess-memory model.

Linux 19 Sep 2026 6 min read

process_madvise Applies Memory Advice Across Process Boundaries

A process can consume memory on behalf of work that is coordinated elsewhere. Linux process_madvise() lets that external coordinator apply selected virtual-memory advice to ranges in the target process without injecting code into it. The target is identified by a pidfd, while the ranges are supplied as an array of struct iovec. That arrangement separates memory-policy decisions from the code that owns the mapping. A runtime manager, service supervisor, or memory controller can request reclaim-oriented or prefetch-oriented treatment for another process, subject to kernel support and permission checks. The system call does not transfer ownership of the mapping, freeze the target, or make its address space stable.

Linux 19 Sep 2026 6 min read

pidfd_getfd Duplicates a Live Descriptor Across Process Boundaries

A process can acquire a new descriptor that refers to the same open file description as a descriptor already held by another process, without asking that target process to send it. Linux provides this operation through pidfd_getfd(). The resulting descriptor is local to the caller, but the kernel object behind it is shared with the target descriptor. That distinction matters because a descriptor number is only an entry in one process’s descriptor table. The open file description carries state such as the current file offset and file status flags. Duplicating across a process boundary therefore transfers access to an existing kernel file instance rather than reopening the pathname or constructing an independent instance.

Linux 19 Sep 2026 5 min read

PID File Descriptors Give Linux a Stable Handle for Process Lifecycle Events

A numeric PID can name one process now and a different process later. Linux PID file descriptors change that boundary: a pidfd is a file descriptor that refers to a task, so process operations can remain attached to the intended kernel object instead of repeating a lookup by numeric PID. This distinction matters in supervisors, service managers, container runtimes, and other software that observes process lifecycles. A PID is useful for naming, but it is not a durable capability. A pidfd can participate in file-descriptor APIs and can be retained across the interval between identifying a process and acting on it.

Tech 19 Sep 2026 7 min read

PCIe Posted Writes Separate CPU Completion from Device Visibility

PCIe Posted Writes Separate CPU Completion from Device Visibility An MMIO store can be complete from the CPU’s point of view while the corresponding write is still moving through the I/O path. PCI and PCIe memory writes are normally posted: the requester does not wait for a completion response for each write. Bridges and interconnect logic can accept the transaction and let the CPU continue before the endpoint has consumed it.