Skip to content

Archive

System Calls

67 articles
Software Engineering 18 Sep 2026 4 min read

timerfd Turns Timer Expiration Counts into Pollable Descriptor State

A Linux timerfd becomes readable when its configured timer has expired. The bytes returned by read(2) are not a timestamp or event record: they encode one unsigned 64-bit integer containing the number of expirations since the previous successful read. Timer state therefore participates in the same readiness machinery as sockets and pipes while retaining timer-specific semantics behind the descriptor boundary. Readiness represents a pending expiration count timerfd_create(2) creates a descriptor associated with a clock, while timerfd_settime(2) arms or disarms its timer. Once at least one expiration is pending, poll(2), select(2), and epoll(7) can report the descriptor as readable.

Linux 18 Sep 2026 4 min read

timerfd Counts Expirations Through Descriptor I/O

A periodic timerfd does not require one userspace wakeup for every timer expiration. If several expirations occur before the descriptor is read, Linux accumulates them and returns the count in one 8-byte integer. That behavior makes timer state fit the same readiness model used for sockets, pipes, and other descriptors. It also gives delayed event loops explicit information about missed periods rather than collapsing several expirations into one notification. Expiration state becomes readable descriptor data timerfd_create() creates a timer object and returns a file descriptor referring to it. The selected clock defines the timer’s time base. Common choices include CLOCK_MONOTONIC, CLOCK_REALTIME, and CLOCK_BOOTTIME.

Software Engineering 18 Sep 2026 3 min read

signalfd Routes Pending Signals Through Descriptor I/O

A Linux signalfd becomes readable when a signal selected by its mask is pending for the reading context. A successful read(2) consumes pending signal state and returns one or more fixed-size signalfd_siginfo records. Signal handling can therefore enter a descriptor-driven event loop without turning asynchronous handlers into the primary dispatch mechanism. The descriptor mask does not block signals The mask passed to signalfd(2) selects signals that the descriptor can accept. It does not modify the calling thread’s signal mask. Normal use separately blocks those signals with sigprocmask(2) or pthread_sigmask(3) so their ordinary dispositions do not run before descriptor consumption.

Linux 18 Sep 2026 5 min read

Seccomp User Notifications Delegate Selected System Calls to a Supervisor

A seccomp filter can stop a selected system call before execution and turn it into a request on a listener file descriptor. The calling thread remains blocked while a userspace supervisor examines the notification and returns a result. This creates a mediation boundary that is narrower than tracing every system call and more dynamic than encoding every decision directly in classic BPF. The mechanism is SECCOMP_RET_USER_NOTIF. A filter returns that action for operations that require external mediation. A filter installed with SECCOMP_FILTER_FLAG_NEW_LISTENER yields a listener file descriptor, and a supervisor uses seccomp notification ioctls on that descriptor.

Software Engineering 18 Sep 2026 6 min read

seccomp User Notification Delegates Selected Syscalls to a Supervisor

seccomp User Notification Delegates Selected Syscalls to a Supervisor A seccomp filter can stop a selected system call before the kernel executes it and emit a notification to a user-space supervisor instead. The target thread remains blocked while the supervisor receives the event and returns a disposition. This behavior turns a filter result into a controlled handoff across the kernel/user-space boundary. The mechanism is SECCOMP_RET_USER_NOTIF. It differs from ordinary seccomp actions because the BPF filter does not finish the decision by itself. A listener file descriptor becomes the coordination point for notification receipt, response delivery, and optional file-descriptor injection.

Linux 18 Sep 2026 4 min read

process_madvise Applies Memory Reclaim Advice Across Process Boundaries

process_madvise() can make one Linux process request memory-management action for virtual-address ranges owned by another process. The target is identified by a pidfd, while an iovec array names the target ranges. This separates memory-policy decisions from the process whose mappings receive the advice. The interface is useful for controllers that already have external knowledge about workload state. A runtime manager can mark inactive memory cold or request page reclamation without injecting code into the managed process. That capability is bounded by permission checks, supported advice values, and partial-progress semantics.

Linux 18 Sep 2026 6 min read

pidfd_getfd Duplicates Another Process File Descriptor into the Caller

A file descriptor number has meaning only inside its process descriptor table, but the kernel object behind that number can be shared across processes. Linux pidfd_getfd() bridges those two scopes: it takes a PID file descriptor plus a descriptor number from the referenced process and installs a duplicate descriptor in the caller. The new descriptor refers to the same open file description as the target descriptor. That last property is the central boundary. pidfd_getfd() does not reopen a pathname, copy bytes, or create an independent file position. It duplicates an existing kernel reference and therefore inherits sharing semantics that can affect both processes.

Linux 18 Sep 2026 5 min read

openat2 Resolve Flags Constrain Path Traversal per Open

A pathname passed to openat2() can be rejected even when the same pathname would resolve successfully through openat(). The difference comes from open_how.resolve: Linux can apply traversal constraints while resolving every component of that single open operation. This changes the boundary around path handling. A directory file descriptor can act as more than a starting point; resolve flags can restrict escapes, symbolic-link traversal, mount crossings, and lookups that require work beyond cached state.

Cybersecurity 18 Sep 2026 5 min read

openat2 Resolution Flags Constrain Path Traversal at the Kernel Boundary

A service can validate a pathname and still open a different object if the namespace changes between validation and use. Symbolic links, mount topology, rename operations, and special procfs links make pathname resolution a kernel operation with state that can change concurrently. Linux openat2() addresses part of this boundary by attaching resolution constraints to the lookup that produces the file descriptor. The security property is narrower than generic path sanitization. openat2() does not declare a pathname safe. It lets a caller ask the kernel to reject specific resolution behavior while the kernel performs the walk.

Software Engineering 18 Sep 2026 5 min read

openat2 Constrains Path Resolution Inside a Directory Boundary

A pathname is not a stable object reference. Between its starting directory and final component, Linux path resolution may follow symbolic links, cross mount points, process .., or encounter special links exposed by pseudo-filesystems. openat2() lets a caller attach constraints to that resolution operation so the kernel can reject a lookup that leaves the intended boundary. The distinction is stronger than checking a normalized string before open(). String validation examines syntax. openat2() can constrain the kernel’s actual traversal while filesystem objects and mount topology participate in the lookup.

Linux 18 Sep 2026 5 min read

mseal Locks Memory Mapping Layout and Permissions

A process can establish a memory mapping with the intended address, size, and protection bits, then later alter that mapping with operations such as munmap(), mprotect(), or mremap(). Linux mseal() adds a one-way state transition: selected virtual memory areas can be sealed so a class of later mapping modifications is rejected by the kernel. The mechanism protects mapping structure rather than the bytes stored in the mapping. A writable sealed mapping remains writable through ordinary stores. Sealing instead constrains operations that could remove the mapping, relocate it, replace it, or change attributes covered by the sealing rules.

Linux 18 Sep 2026 6 min read

membarrier Moves Memory-Ordering Cost to an Infrequent Coordination Path

A full hardware memory barrier in a frequently executed path can impose a cost on every operation, even when cross-thread coordination happens only occasionally. Linux membarrier() supports a different placement of that cost: a rare coordination path can request an ordering event across a defined set of threads while a frequent path may need only compiler-level ordering. This is not a general replacement for atomics, locks, or the memory model of a programming language. It is a Linux-specific synchronization primitive for designs whose correctness already has a precise pairing between a frequent path and an infrequent coordination path.

Software Engineering 18 Sep 2026 6 min read

Linux userfaultfd Moves Selected Page Fault Handling into User Space

A page fault normally crosses from a process into the kernel and returns only after the kernel has resolved the virtual-memory condition or delivered an error. Linux userfaultfd can insert a user-space component into that path for explicitly registered address ranges. The kernel reports selected faults through a file descriptor, blocks the faulting execution context when the mode requires it, and accepts an ioctl that resolves the fault. This is a Linux virtual-memory interface, not a C or POSIX memory guarantee. Its behavior depends on negotiated kernel features, the registered range, its mapping type, and the registration mode.

Software Engineering 18 Sep 2026 7 min read

Linux splice Makes Pipe Capacity Part of Data-Transfer Semantics

Linux splice() can transfer bytes between file descriptors without routing those bytes through a user-space buffer, but the interface is not a generic descriptor-to-descriptor copy primitive. At least one endpoint must be a pipe. That requirement makes pipe state part of the transfer contract: capacity, readable data, writer presence, blocking mode, and partial progress can all affect an otherwise straightforward data path. The useful boundary is therefore not simply “kernel copy versus user copy.” splice() changes the shape of ownership and flow control. Application code stops owning an intermediate byte array, while it still owns the control loop that accounts for bytes transferred, handles readiness, and preserves offset semantics.

Software Engineering 18 Sep 2026 5 min read

Linux signalfd Converts Selected Signals into Descriptor Reads

A signal included in a signalfd mask can become readable state on a file descriptor instead of invoking an asynchronous signal handler, provided that the signal is blocked from ordinary delivery in the relevant thread. This changes the interface boundary: signal arrival can participate in the same descriptor-oriented event loop as sockets, pipes, and other pollable objects. The mechanism does not replace Linux signal semantics. Signal generation, process and thread signal masks, pending state, standard-signal coalescing, real-time signal queuing, and delivery rules still apply. signalfd changes the consumption interface for signals selected by its mask.

Software Engineering 18 Sep 2026 7 min read

Linux signalfd Converts Pending Signals into Descriptor Reads

Linux signalfd gives selected signals a descriptor-oriented consumption path. Instead of transferring control into an asynchronous handler, a process can block those signals, associate them with a signalfd object, and consume pending instances through read(). The descriptor can also participate in poll(), select(), and epoll, placing signal reception beside sockets, timers, and other readiness sources. This interface is Linux-specific. The signal mask, pending-signal rules, and descriptor operations come from Linux and POSIX signal semantics where applicable; they are not properties of the C language itself.

Software Engineering 18 Sep 2026 6 min read

Linux renameat2 Makes Path-Replacement Policy Atomic

A pathname rename changes directory entries while open file descriptors continue to refer to the same underlying objects. Linux renameat2() adds policy to that namespace update: a caller can reject replacement, exchange two existing names, or request a whiteout for union-filesystem operation. These policies are executed as part of the rename operation rather than as checks performed separately in userspace. The interface is Linux-specific. A zero flags argument gives renameat() behavior, while nonzero flags add Linux semantics that also depend on support from the mounted filesystem.

Software Engineering 18 Sep 2026 5 min read

Linux pidfd Binds Process Operations to Stable Kernel Identity

A numeric PID names a process through a namespace lookup. That number can later be reused after the process exits and is reaped. Linux PID file descriptors change the boundary: a pidfd is a file descriptor referring to a particular task, so later operations can target that reference instead of resolving the numeric PID again. This is Linux-specific process-management behavior. It is not a property of POSIX process identifiers or of the C language.

Software Engineering 18 Sep 2026 6 min read

Linux openat2 Constrains Path Resolution Inside a Directory Boundary

A pathname passed to openat2() can be resolved relative to a directory file descriptor while the kernel enforces restrictions on the resolution process itself. That distinction matters when a process accepts path components from a less-trusted source. A string check can inspect the pathname text, but it cannot by itself freeze the filesystem namespace while lookup proceeds. openat2() places the policy beside the lookup. Its struct open_how separates ordinary open flags from resolve flags that constrain traversal. The resulting boundary is about resolution semantics, not merely the spelling of a path.

Software Engineering 18 Sep 2026 5 min read

Linux eventfd Represents Counter State Through Descriptor Readiness

An eventfd object stores an unsigned 64-bit counter in the kernel and exposes that state through a file descriptor. Writes add to the counter under defined bounds; reads consume counter state; readiness interfaces expose whether an operation can proceed without blocking. The result is a compact synchronization boundary that fits descriptor-oriented event loops without turning the counter into a byte stream. The interface is Linux-specific. Its guarantees come from the eventfd system-call contract and kernel descriptor semantics, not from the C language or POSIX.

Software Engineering 18 Sep 2026 5 min read

Linux close_range Makes Descriptor Cleanup a Table Operation

A process preparing for execve() often needs a simple invariant: descriptors above a small allowlist must not survive into the new program. Closing descriptor numbers one by one turns that invariant into an enumeration problem. Linux close_range() expresses it directly as an operation over an inclusive interval of the calling task’s file descriptor table. The interface is Linux-specific. Its behavior belongs to Linux file-table and system-call semantics, not to the C language or a portable POSIX guarantee.

Linux 18 Sep 2026 6 min read

Landlock Rulesets Add Process-Local Access Control

A Linux process can voluntarily remove access that its UID, capabilities, mount namespace, and other security layers would otherwise permit. Landlock implements this as a stackable Linux Security Module: a process creates a ruleset, adds allowed objects, then places itself in a Landlock domain. The resulting policy is an additional restriction. It does not grant access denied by DAC, ACLs, SELinux, AppArmor, mount permissions, or another active control. Once enforced, the Landlock layer cannot be removed from that thread; later Landlock domains can only add restrictions.

Linux 18 Sep 2026 5 min read

eventfd Aggregates Notifications in a Kernel Counter

An eventfd can absorb several notification writes before userspace services the descriptor. The kernel stores those writes in a 64-bit counter, so readiness represents pending counter state rather than a queue containing one record per notification. That distinction matters in event loops. A producer can add values while a consumer is occupied, and the next read can collapse accumulated state into one result. With EFD_SEMAPHORE, the same object exposes a different consumption rule without changing its readiness model.

Linux 18 Sep 2026 6 min read

close_range Controls Descriptor Inheritance Across exec Boundaries

A process preparing to call execve() may have hundreds or thousands of open file descriptors, while the new program should inherit only a small selected set. Closing descriptors one by one creates both bookkeeping cost and a concurrency problem: another thread can allocate a descriptor while the cleanup loop is still running. Linux close_range() moves that operation to the descriptor-table boundary. A caller specifies an inclusive numeric range and asks the kernel either to close descriptors in that range or mark them close-on-exec. With CLOSE_RANGE_UNSHARE, the caller can first detach its descriptor table from threads or processes that share it.