Skip to content

Archive

System Calls

67 articles
Linux 23 Sep 2026 4 min read

TFD_TIMER_CANCEL_ON_SET Turns Realtime Clock Jumps into ECANCELED

An absolute timer tied to CLOCK_REALTIME has a dependency that a monotonic deadline does not: an administrator, synchronization service, or privileged process can move the wall clock discontinuously while the timer is armed. Linux timerfd can expose that event explicitly. With TFD_TIMER_ABSTIME | TFD_TIMER_CANCEL_ON_SET, a qualifying clock jump causes a current or later read() on the timer descriptor to fail with ECANCELED. This behavior separates two events that would otherwise be easy to conflate: reaching a scheduled wall-clock instant and invalidating the clock basis used to schedule it.

Linux 23 Sep 2026 4 min read

pidfd_getfd Duplicates a Target Descriptor into the Calling Process

pidfd_getfd Duplicates a Target Descriptor into the Calling Process pidfd_getfd() can place a duplicate of another process’s open file descriptor into the caller’s descriptor table. The returned descriptor is local to the caller, but it refers to the same open file description as the selected descriptor in the target process. That distinction matters. This operation does not reopen a pathname, reconstruct a socket, or create an independent file position. It duplicates an existing kernel reference across a process boundary.

Linux 23 Sep 2026 5 min read

membarrier Private Expedited Orders Userspace Memory Across Threads

Linux membarrier() can move part of a synchronization cost from a frequently executed path to a less frequent coordination path. With MEMBARRIER_CMD_PRIVATE_EXPEDITED, one thread asks the kernel to establish a memory-ordering point across the running threads in the same process. The caller pays for the system call when coordination is needed instead of requiring every fast-path execution to carry an explicit hardware memory barrier. This is narrower than a general thread rendezvous. membarrier() does not run an application callback on sibling threads, does not wait for application-level acknowledgements, and does not turn ordinary data races into valid synchronization. Its contract concerns ordering of userspace memory accesses around the barrier.

Linux 23 Sep 2026 5 min read

close_range with UNSHARE Separates File Descriptor Cleanup from Peer Threads

A multithreaded Linux process can share one file descriptor table across its threads. That arrangement is convenient during normal execution: a descriptor opened by one thread becomes available to its peers. It becomes less convenient when one thread is preparing a restricted execution context and wants to discard a broad descriptor range without racing with peers that can still allocate descriptors. close_range() provides a range operation for this boundary. With CLOSE_RANGE_UNSHARE, the kernel first separates the caller from the shared descriptor table and then applies the requested closure to the caller’s resulting table. The operation is close to combining unshare(CLONE_FILES) with a range close, but the kernel can perform less work in common cases.

Linux 22 Sep 2026 5 min read

TFD_TIMER_CANCEL_ON_SET Reports Realtime Clock Discontinuities

A timerfd armed against wall-clock time can cross a discontinuous clock adjustment before its deadline. With TFD_TIMER_CANCEL_ON_SET, Linux exposes that event through the descriptor: a subsequent read() fails with ECANCELED rather than presenting the clock jump as an ordinary timer expiration. The flag is deliberately narrow. It applies only when timerfd_settime() arms an absolute timer on CLOCK_REALTIME or CLOCK_REALTIME_ALARM, and it changes the handling of discontinuous changes to those clocks.

Linux 22 Sep 2026 6 min read

signalfd Turns Blocked Signals into Readable Events

Linux signals normally interrupt a thread through asynchronous delivery. signalfd() offers a different boundary for selected signals: keep them blocked in the relevant threads, then consume pending instances by reading a file descriptor. The signal mechanism itself does not become a byte stream. The kernel still maintains signal state and normal process- or thread-directed delivery rules. signalfd adds a file-descriptor interface for accepting signals from a configured set. Blocking and the descriptor mask are separate state A signalfd has a signal-set mask that selects which signals it can accept. That mask does not block those signals for a thread. Normal use therefore pairs descriptor creation with a signal mask operation.

Linux 22 Sep 2026 5 min read

EFD_SEMAPHORE Makes eventfd Reads Consume One Counter Unit

An eventfd normally turns its entire nonzero counter into one read result and resets the counter to zero. Creating it with EFD_SEMAPHORE changes only the read side: each successful read returns the 64-bit value 1 and subtracts one from the kernel-maintained counter. That difference lets several units accumulated by writers remain separately consumable. The object is still an eventfd, with the same counter, write rules, descriptor lifetime, and readiness integration.

Linux 21 Sep 2026 5 min read

openat2 Resolve Flags Constrain Path Traversal During Lookup

A pathname passed to openat() is resolved by the kernel, but the caller has limited control over traversal through intermediate components. Linux openat2() adds a resolve field that applies constraints to the complete lookup operation. The restriction is evaluated while components are traversed, rather than by validating a pathname in userspace and opening it later. That distinction matters when path components can change concurrently. A userspace sequence that checks a path and then opens it creates separate observations of mutable filesystem state. openat2() places the selected lookup policy in the same system call that returns the file descriptor.

Linux 21 Sep 2026 6 min read

memfd File Seals Freeze Shared Memory State

A memfd_create() file starts as a mutable anonymous file. It can be resized, written, and mapped much like a regular file, while its storage remains volatile and disappears after the last reference is released. File sealing adds a different phase to that lifecycle: after data has been populated, the kernel can permanently reject selected classes of later modification. That transition is useful when one process prepares bytes and then hands the same file description to another process. The receiver can inspect the seals attached to the inode instead of relying only on a convention that the sender will stop changing the object.

Linux 21 Sep 2026 6 min read

io_uring Multishot Accept Keeps One Accept Request Active

A normal accept request has a one-to-one shape: one submitted operation eventually yields one completion. io_uring multishot accept changes that relationship. A single accept SQE can remain active across multiple incoming connections and emit a separate completion queue entry for each accepted socket. The request is persistent, but not permanent. Each CQE carries enough state for the application to tell whether the original request can produce another completion. That boundary matters because a server that treats every successful CQE as proof that accept is still armed can silently stop accepting after the multishot request terminates.

Linux 21 Sep 2026 4 min read

futex_waitv Blocks on Multiple Futex Words in One Wait

A traditional futex wait names one futex word. That maps cleanly to a mutex or condition whose blocking state is represented by one shared 32-bit value. Some synchronization designs instead need a thread to sleep until any member of several independent states changes. futex_waitv() provides that vector wait. The caller supplies an array of wait descriptors, each containing a futex address and an expected value. The kernel checks the vector and blocks only while every entry still matches its expected state.

Linux 20 Sep 2026 5 min read

pidfd_getfd Duplicates a Live File Descriptor Across Processes

A process can acquire a usable duplicate of a file descriptor that is already open in another process without asking that process to send it over a UNIX domain socket. Linux pidfd_getfd() performs that transfer through a PID file descriptor, subject to a ptrace access check. The returned descriptor is new in the caller, but the kernel object behind it is not independent. It refers to the same open file description as the target descriptor. That distinction controls offset sharing, file status flags, and operations on the underlying object.

Linux 19 Sep 2026 5 min read

userfaultfd Moves Selected Page-Fault Resolution into User Space

A memory access normally enters the kernel page-fault path and completes without an application choosing the page contents at that instant. Linux userfaultfd changes that boundary for registered virtual address ranges: selected faults become events on a file descriptor, and a user-space manager can supply or activate the page before the faulting thread continues. The mechanism does not replace the process page tables with a user-space data structure. The kernel still owns page-table state and performs the final mapping operation. User space gains control over specific fault classes and the timing of their resolution.

Linux 19 Sep 2026 5 min read

timerfd Reads Count Periodic Expirations

A periodic timer can expire several times before a busy event loop gets CPU time again. Linux timerfd does not compress that delay into a bare “timer fired” notification. A successful read() returns an unsigned 64-bit count of expirations accumulated since the timer was armed or since the preceding successful read. That counter changes the semantics of delayed timer handling. Readiness says at least one expiration is pending; the value read from the descriptor says how many periods elapsed.

Linux 19 Sep 2026 4 min read

signalfd Converts Linux Signals into File-Descriptor Events

A conventional POSIX signal can interrupt a thread at almost any instruction boundary and transfer control to a signal handler. That asynchronous control flow imposes strict limits on handler code and complicates programs whose main control plane already runs through epoll, poll, or select. Linux signalfd() provides a different delivery interface. A process blocks selected signals in the normal signal mask, creates a signalfd for that set, and receives pending signals by reading structured signalfd_siginfo records from the descriptor. The signal remains a signal at the kernel interface; only its consumption moves into ordinary file-descriptor I/O.

Linux 19 Sep 2026 5 min read

seccomp User Notifications Move Selected System Calls to a Supervisor

seccomp User Notifications Move Selected System Calls to a Supervisor A seccomp filter can do more than allow a system call or reject it in the kernel. When a filter returns SECCOMP_RET_USER_NOTIF, Linux can suspend the calling thread and deliver a description of that attempted system call to a user-space supervisor. The mechanism creates an interposition boundary around selected calls. It is useful when a less-privileged process needs an operation mediated by another process, such as a container manager handling a call that the container cannot perform directly. The boundary is deliberately narrower than a general security-policy engine: notification state can race with mutable target memory, and the kernel documentation warns against treating the supervisor’s inspection as an authorization primitive.

Software Engineering 19 Sep 2026 6 min read

seccomp User Notification Moves Selected Syscall Decisions to a Broker

A seccomp filter can do more than allow or reject a system call immediately. With user notification, a matching call can be suspended while another process receives a structured request on a listener file descriptor and decides what result the blocked thread receives. The mechanism turns selected syscall decisions into a brokered interface without moving the entire syscall implementation into user space. The boundary is precise but narrower than a general interposition layer. The kernel still owns syscall dispatch, task state, descriptor tables, and validation performed by kernel code. The broker receives metadata and can return a value, an error, or in supported cases request continued execution of the original syscall. Correct designs account for mutable target memory, notification lifetime, and the fact that a policy decision is not automatically a transaction over process state.

Software Engineering 19 Sep 2026 5 min read

process_vm_readv and process_vm_writev Transfer Memory Across Process Boundaries

process_vm_readv() and process_vm_writev() let one Linux process copy bytes directly between its address space and another process’s address space. The calls operate on vectors of local and remote memory ranges, but a successful process lookup does not make remote memory stable. Mapping changes, page accessibility, permissions, and concurrent mutation remain separate parts of the contract. These interfaces are Linux-specific system calls. They do not define C object lifetime, synchronization, or a portable interprocess-memory model.

Linux 19 Sep 2026 6 min read

openat2 Constrains Path Resolution at the Kernel Boundary

A pathname that begins inside a trusted directory can resolve somewhere else before open() returns. Parent components, symbolic links, magic links, mount points, and concurrent namespace changes all participate in Linux pathname lookup. Checking a string before opening it therefore does not establish where the kernel will finish resolution. Linux openat2() places restrictions inside the lookup operation itself. A caller supplies a directory file descriptor, ordinary open flags, and a resolve policy in struct open_how. The kernel then applies those constraints while walking every relevant path component. This moves a security boundary from pre-validation of pathname text into the operation that actually resolves the pathname.

Linux 19 Sep 2026 5 min read

openat2 Constrains Linux Path Resolution at the Open Boundary

A pathname can change meaning while a process is resolving it. Directory renames, symbolic links, mount points, and .. components can redirect lookup away from the directory a program intended to treat as its boundary. Linux openat2() attaches resolution policy to the lookup itself. Its struct open_how contains a resolve bit mask, so the kernel can reject a path when resolution violates a caller-selected constraint instead of relying only on checks performed before open().

Software Engineering 19 Sep 2026 6 min read

Linux pidfds Bind Process Operations to Stable Kernel References

A numeric process ID names a process only while that PID remains assigned to it. After process exit and reaping, Linux may reuse the number for another process. Code that observes a PID, performs unrelated work, then acts on that number can therefore cross a lifetime boundary that the integer itself does not encode. Linux pidfds provide a file-descriptor reference to a process so later operations can target the referenced process object rather than repeat a numeric PID lookup.

Linux 19 Sep 2026 5 min read

eventfd Turns Kernel Notifications into Pollable Counters

A Linux process can signal work through a file descriptor without moving a byte stream between producer and consumer. eventfd() creates a kernel-maintained 64-bit counter whose readiness can be observed by poll(), select(), or epoll. A write adds to the counter; a read consumes its accumulated state according to the descriptor mode. That shape makes eventfd different from a pipe. A pipe preserves a sequence of bytes. An eventfd preserves counter state. When the application needs a wakeup edge plus a compact amount of accumulated state, that distinction removes buffering and framing that a byte stream would otherwise require.

Linux 19 Sep 2026 5 min read

close_range Makes File-Descriptor Cleanup a Single Linux Operation

A process preparing to execute another program often needs a simple boundary: descriptors 0, 1, and 2 remain available, while every higher descriptor must disappear. Repeating close() over a guessed numeric limit or enumerating /proc/self/fd turns that boundary into a userspace scan. Linux close_range() expresses the interval directly. The kernel applies one operation to every open file descriptor from first through last, inclusive. With flags, the same interface can isolate a shared descriptor table or mark the interval close-on-exec instead of closing it immediately.

Linux 18 Sep 2026 5 min read

userfaultfd Turns Page Faults into Userspace Events

A thread can fault on a virtual address and remain blocked while another userspace thread or process decides what page state should make that access continue. userfaultfd provides this boundary by turning selected page faults into messages on a file descriptor and pairing those messages with ioctls that resolve the fault. The mechanism does not replace the kernel page-fault machinery. It inserts userspace control at registered ranges and fault classes, while the kernel still owns page tables, fault blocking, and the transition that makes the page usable again.