Skip to content

Archive

System Calls

67 articles
Cybersecurity 17 Sep 2026 6 min read

Seccomp User Notification Moves Selected Syscall Decisions to a Supervisor

A confined process issues a system call that its ordinary seccomp policy cannot safely reduce to a static allow-or-deny decision. The arguments may refer to mutable process memory, or the operation may need privileged work performed outside the confined process. Returning a fixed errno is too restrictive, while permitting the call directly gives the target more authority than the deployment intends. Linux seccomp user notification creates a mediation path for this case. A filter can return SECCOMP_RET_USER_NOTIF, causing the kernel to block the triggering task and emit a request on a listener file descriptor. A userspace supervisor receives the request and later supplies a result. This mechanism changes where a selected syscall decision is made, but it does not turn seccomp into a general reference monitor without additional policy and race controls.

Software Engineering 17 Sep 2026 6 min read

Seccomp User Notification Delegates Selected System Calls to a Supervisor

A system call selected by a seccomp filter can stop before kernel execution and appear as a request on a notification file descriptor. With SECCOMP_RET_USER_NOTIF, Linux turns that call into a coordination point between the blocked target and a userspace supervisor. The supervisor can emulate a result, inject a file descriptor for suitable operations, or permit the kernel to continue the original call. This mechanism is deliberately narrower than a general userspace security policy engine. Its strongest boundary is the kernel-mediated suspension and response protocol. Data reached through target-memory pointers can still change around a supervisor’s inspection, and a response that continues the original call re-enters ordinary kernel execution with that race still relevant.

Software Engineering 17 Sep 2026 4 min read

renameat2 RENAME_EXCHANGE Swaps Two Paths in One Filesystem Operation

renameat2() with RENAME_EXCHANGE changes two existing directory entries as one atomic rename operation. Before the call, each pathname reaches its original object; after a successful call, each pathname reaches the object formerly named by the other path. There is no successful intermediate state in which one of the two names has merely been removed or overwritten. That property is distinct from ordinary rename(). A conventional rename can atomically replace a destination, but replacement discards the destination name from the namespace. Exchange preserves both named objects and swaps their positions.

Linux 17 Sep 2026 5 min read

rename Replaces a Directory Entry Atomically on Linux

A successful rename() can replace an existing destination pathname without exposing an intermediate state in which that destination name is missing. Processes resolving the destination observe either the old directory entry or the replacement, subject to filesystem and mount constraints. That atomic namespace transition is narrower than several properties often associated with file replacement. It does not make prior writes durable, does not force directory metadata to stable storage, and does not invalidate file descriptors that already refer to the replaced file.

Linux 17 Sep 2026 4 min read

O_CLOEXEC Closes Descriptors Atomically Across exec

A file descriptor created without close-on-exec state can escape into a newly executed program during a narrow concurrency window. In a multithreaded Linux process, setting FD_CLOEXEC with a later fcntl() call leaves that window open between descriptor creation and the flag update. O_CLOEXEC removes the split operation. The kernel creates the descriptor with its close-on-exec flag already set, so another thread cannot observe an intermediate state in which the descriptor exists but remains inheritable across a successful execve().

Linux 17 Sep 2026 5 min read

O_APPEND Couples End Positioning with Each Write

O_APPEND changes a write from two separable actions into one coupled operation: Linux positions the open file description at the current end of the file and performs the write as a single atomic step. That property matters when multiple writers target one regular file. A sequence built from lseek(fd, 0, SEEK_END) followed by write(fd, ...) does not carry the same append semantics because another writer can change the file between those two system calls.

Software Engineering 17 Sep 2026 4 min read

memfd Seals Turn Shared Files into Monotonic Objects

A Linux memfd can begin as a writable anonymous file and later acquire restrictions that cannot be removed. The restrictions belong to the inode, so transferring or duplicating a descriptor does not create a less restricted view. Once a seal is added successfully, every descriptor referring to that inode is subject to it. This makes sealing different from descriptor access modes. A descriptor can carry local flags, while a seal changes the mutation boundary of the shared file object itself.

Software Engineering 17 Sep 2026 6 min read

Linux seccomp Notification Splits Syscall Entry From Supervisor Response

A seccomp filter can stop a task at syscall entry and turn that event into a message for another process. With SECCOMP_RET_USER_NOTIF, the kernel does not immediately execute the selected syscall. It creates a notification for a listener, blocks the calling task, and waits for a response that can supply a return value, inject a file descriptor, or permit the syscall to continue. That boundary is narrower than general syscall emulation. The notification carries syscall metadata and register argument values, while memory referenced by pointer arguments remains in the target process. The target can also disappear or have its notification invalidated while a supervisor is making a decision. Those properties make identity, memory ownership, and response timing part of the interface contract.

Software Engineering 17 Sep 2026 6 min read

Linux openat2 Makes Path Resolution Policy Part of the Open Operation

A pathname is not an object reference. It is an instruction for traversing a mutable namespace, and another task can alter directory entries, symbolic links, or mounts while that traversal is relevant to an application. Linux openat2() addresses this boundary by placing path-resolution constraints in the same kernel operation that returns the file descriptor. That placement matters when a program accepts a pathname but intends to confine resolution to a directory tree. A user-space sequence that inspects components and later calls open() separates validation from use. openat2() can instead make selected traversal rules part of the lookup itself.

Linux 17 Sep 2026 4 min read

Duplicated File Descriptors Share an Open File Description

Two file descriptor numbers can move the same file offset. On Linux, this occurs when both descriptors refer to one open file description, as happens after dup() and across inherited descriptors after fork(). The distinction matters because a file descriptor is a process-visible integer, while the open file description is the kernel object that carries state for an open instance of a file. Treating those layers as interchangeable can produce offset interference, status-flag changes that cross descriptor boundaries, and surprising behavior after process creation.

Cybersecurity 17 Sep 2026 5 min read

close_range with UNSHARE Detaches Descriptor Tables Before Bulk Closure

A multithreaded Linux process can reach an awkward boundary just before execve(): one thread wants to discard every file descriptor above standard input, output, and error, while another thread can still create descriptors in the same table. A loop of close() calls treats descriptor numbers individually, but it does not by itself change the fact that the table is shared. close_range() with CLOSE_RANGE_UNSHARE addresses that specific race. The kernel first gives the caller a file descriptor table that is no longer shared with the other users of the old table, then applies the requested bulk closure to the caller’s table. The security property is about table ownership during cleanup, not merely fewer system calls.

Linux 07 Sep 2026 12 min read

Understand Sparse Files on Linux Without Wasting Disk Space

A file can report a size of 100 GiB without consuming 100 GiB of disk blocks. That sounds contradictory until you separate two ideas that ordinary file APIs often present together: a file’s logical size and the storage that the filesystem has actually allocated for it. Linux filesystems can represent long ranges of unwritten bytes as holes. Reading a hole returns zero bytes, but the filesystem does not need to store one physical zero byte for every logical byte in that range. A regular file that contains holes is called a sparse file.

Linux 06 Sep 2026 13 min read

Choose the Right Advisory File Lock on Linux

Two processes can open the same file and both write to it successfully. That is often exactly what Unix applications need, but sometimes the programs are supposed to coordinate: only one worker should update a state file, several readers may share a resource, or a process must avoid changing a byte range while another process is using it. Linux offers several advisory file-locking mechanisms. The confusing part is not how to request a lock. The confusing part is what owns the lock and when that lock disappears.

Linux 05 Sep 2026 11 min read

Wake Linux Event Loops from Other Threads with eventfd

A file-descriptor event loop can wait efficiently for sockets, pipes, timers, and other kernel objects. A common problem appears when work originates somewhere that is not already represented by a file descriptor: another thread changes shared state and needs the loop to wake immediately. Polling shared state on a timer adds latency or wastes wakeups. A condition variable can wake a thread, but it cannot be placed directly in the same poll() or epoll wait set as a socket. A pipe can bridge the two models, but using a pipe only as a wakeup signal means maintaining a read end, a write end, and byte-buffer semantics that the application may not actually need.

Linux 05 Sep 2026 10 min read

Move Data Between Linux File Descriptors with splice()

A conventional file-copy loop reads bytes into a user-space buffer and then writes those bytes somewhere else. That pattern is portable and easy to understand, but sometimes the program does not need to inspect or transform the data at all. It only needs to move bytes from one file descriptor to another. On Linux, splice() can handle that case differently. It transfers data between file descriptors while keeping the transferred data out of a user-space buffer. At least one endpoint must be a pipe, so a pipe can act as the kernel-side bridge between a source and a destination.

Linux 05 Sep 2026 9 min read

Create Sealable In-Memory Files on Linux with memfd_create

Applications often need a temporary chunk of data that behaves like a file without needing a persistent pathname. A process may build a configuration snapshot, compiled artifact, or serialized message, then map it into memory or pass it to another process. A regular temporary file can do that, but it introduces filesystem naming, cleanup, permissions, and lifetime concerns. An anonymous mmap() avoids the pathname, but it does not produce an ordinary file descriptor that can be passed to APIs expecting file-backed data.

Linux 04 Sep 2026 11 min read

Integrate Timers into Linux Event Loops with timerfd

Event loops work best when unrelated kinds of work have one common waiting mechanism. Sockets become readable. Pipes become writable. A child-process descriptor or signal descriptor can become ready. Timers are often the awkward exception. A program can call sleep() or nanosleep(), but that blocks the thread instead of letting it wait for I/O. It can pass a timeout to poll() or epoll_wait(), but one timeout becomes difficult to manage when the program has several independent deadlines. Traditional POSIX timers can deliver signals, which introduces a second asynchronous control path.

Linux 04 Sep 2026 9 min read

Handle Linux Signals in Event Loops with signalfd

Unix signals are asynchronous by design: a signal can interrupt a program between ordinary instructions and transfer control to a signal handler. That model is useful, but it creates an awkward boundary for event-driven programs. A network server may already spend most of its time inside poll(), epoll_wait(), or another readiness API. Its sockets, pipes, and timers appear as file-descriptor events, while SIGTERM and SIGHUP arrive through a separate execution path with much stricter rules about what code may safely run.

Linux 04 Sep 2026 8 min read

Avoid PID Reuse Races on Linux with pidfds

A process ID looks like an identity, but it is really a reusable number. That distinction matters in long-running supervisors, job managers, test harnesses, and other programs that observe a process and then act on it later. Between those two operations, the original process can exit and Linux can eventually reuse the same PID for an unrelated process. Traditional PID-based code can therefore have a time-of-check/time-of-use race: check PID 4242 -> original process exits -> PID 4242 is reused -> signal PID 4242 Linux PID file descriptors, usually called pidfds, provide another model. Instead of repeatedly identifying a task by a reusable integer, a program obtains a file descriptor that refers to a particular process and can use that descriptor with pidfd-aware APIs.