Skip to content

Topic archive

Software Engineering

Software Engineering covers developer tooling, architecture, maintainability, engineering workflows, protocols, and practices that improve how software is designed and built.

460 articles
Software Engineering 18 Sep 2026 5 min read

Linux pidfd Binds Process Operations to Stable Kernel Identity

A numeric PID names a process through a namespace lookup. That number can later be reused after the process exits and is reaped. Linux PID file descriptors change the boundary: a pidfd is a file descriptor referring to a particular task, so later operations can target that reference instead of resolving the numeric PID again. This is Linux-specific process-management behavior. It is not a property of POSIX process identifiers or of the C language.

Software Engineering 18 Sep 2026 6 min read

Linux openat2 Constrains Path Resolution Inside a Directory Boundary

A pathname passed to openat2() can be resolved relative to a directory file descriptor while the kernel enforces restrictions on the resolution process itself. That distinction matters when a process accepts path components from a less-trusted source. A string check can inspect the pathname text, but it cannot by itself freeze the filesystem namespace while lookup proceeds. openat2() places the policy beside the lookup. Its struct open_how separates ordinary open flags from resolve flags that constrain traversal. The resulting boundary is about resolution semantics, not merely the spelling of a path.

Software Engineering 18 Sep 2026 5 min read

Linux O_PATH Separates Object Reference From I/O Authority

open() usually combines two effects: pathname resolution selects a filesystem object, then the returned file descriptor carries an access mode for data I/O. Linux O_PATH splits those effects. A successful open(path, O_PATH) returns a descriptor that refers to the selected object while ordinary read() and write() through that descriptor are not permitted. That split is useful anywhere a process needs a durable kernel reference for later metadata or pathname-relative operations without opening the object for data transfer. It also changes race analysis: later operations can start from the descriptor rather than resolving the original pathname again.

Software Engineering 18 Sep 2026 4 min read

Linux memfd Seals Turn Mutable Memory Files into Enforced State Transitions

A file created by memfd_create() can begin as mutable storage and later acquire kernel-enforced restrictions that apply to the underlying file rather than to one descriptor. With MFD_ALLOW_SEALING, a process can add seals through fcntl(F_ADD_SEALS) and make selected mutations unavailable to every holder of that file. This creates a state transition that ordinary descriptor permissions do not express. A producer can populate bytes, fix the file’s size, and then publish the descriptor with restrictions that remain attached even after the descriptor crosses a process boundary.

Software Engineering 18 Sep 2026 6 min read

Linux memfd Seals Convert Mutable Anonymous Files into Restricted Capabilities

A file descriptor returned by memfd_create() can begin as a writable, resizable anonymous file and later become an object whose permitted mutation operations have been permanently reduced. Linux implements that transition with file seals. The mechanism is attached to the underlying file rather than to one descriptor, so passing a duplicate descriptor across a process boundary does not create an independent sealing state. This property makes sealing more than a convenience around temporary storage. It changes the authority carried by every descriptor that refers to the same memfd object. The transition is monotonic: seals can be added, but they cannot be removed.

Software Engineering 18 Sep 2026 5 min read

Linux inotify Reports Directory-Entry Events, Not Durable Path Identity

An inotify watch does not make a pathname a durable identifier. Linux attaches a watch to a filesystem object selected when inotify_add_watch() succeeds, then emits records describing activity associated with watched objects and directory entries. Names can move, objects can disappear, and event delivery can lose detail when the queue overflows. That boundary matters for file synchronizers, configuration reloaders, indexers, and service supervisors. An event stream can signal that local filesystem state changed, but reconstructing authoritative state still depends on filesystem operations performed after the event.

Software Engineering 18 Sep 2026 5 min read

Linux eventfd Represents Counter State Through Descriptor Readiness

An eventfd object stores an unsigned 64-bit counter in the kernel and exposes that state through a file descriptor. Writes add to the counter under defined bounds; reads consume counter state; readiness interfaces expose whether an operation can proceed without blocking. The result is a compact synchronization boundary that fits descriptor-oriented event loops without turning the counter into a byte stream. The interface is Linux-specific. Its guarantees come from the eventfd system-call contract and kernel descriptor semantics, not from the C language or POSIX.

Software Engineering 18 Sep 2026 8 min read

Linux copy_file_range Separates Copy Semantics From Data Movement

copy_file_range() asks Linux to copy bytes between regular files without requiring the application to shuttle those bytes through a user-space buffer. The call defines a byte-range operation, but it does not prescribe the physical transfer mechanism. A filesystem can perform ordinary data movement, use a copy-on-write sharing mechanism such as reflink, or employ another supported acceleration path while preserving the visible file contents required by the operation. That separation is the central API boundary. Applications specify source and destination ranges and observe the returned byte count. The kernel and filesystem retain latitude over the mechanism used to realize the copy.

Software Engineering 18 Sep 2026 8 min read

Linux close_range Makes Descriptor-Table Cleanup a Range Operation

A process preparing to execute another program often needs a simple descriptor invariant: standard input, output, and error remain available, while unrelated descriptors do not cross the execution boundary. Closing descriptors one at a time can turn that invariant into an enumeration problem. Linux close_range() instead applies an operation to an inclusive numeric interval in the calling task’s file-descriptor table. The interface is small, but its semantics reach into descriptor-table sharing, execve() inheritance, concurrent descriptor allocation, and privilege transitions. The flags select more than implementation strategy: they determine whether descriptors disappear immediately, become close-on-exec, or are first separated from a table shared with other tasks.

Software Engineering 18 Sep 2026 5 min read

Linux close_range Makes Descriptor Cleanup a Table Operation

A process preparing for execve() often needs a simple invariant: descriptors above a small allowlist must not survive into the new program. Closing descriptor numbers one by one turns that invariant into an enumeration problem. Linux close_range() expresses it directly as an operation over an inclusive interval of the calling task’s file descriptor table. The interface is Linux-specific. Its behavior belongs to Linux file-table and system-call semantics, not to the C language or a portable POSIX guarantee.

Software Engineering 18 Sep 2026 4 min read

Landlock Handled Rights Define a Deny-by-Default Sandbox Boundary

A Landlock ruleset does not implicitly deny every operation known to the running kernel. It first declares which access rights it handles. Once the ruleset is enforced, those handled actions are denied by default unless a matching rule grants them. That explicit boundary is central to Landlock compatibility. User space can restrict rights it knows and has tested while a newer kernel may expose additional rights that an older binary never named.

Software Engineering 18 Sep 2026 4 min read

io_uring Links Serialize Dependent Requests

Two adjacent io_uring submission queue entries are normally independent requests. Setting IOSQE_IO_LINK on the first changes that relationship: the next request does not start before the linked request completes. Repeating the flag forms an ordered chain inside one submission batch. The ordering property is narrower than global queue serialization. Requests outside the chain can still run independently, and separate chains can overlap. A link therefore expresses dependency between specific SQEs rather than imposing a barrier on the entire ring.

Software Engineering 18 Sep 2026 4 min read

eventfd Turns Counter State into Descriptor Readiness

An eventfd descriptor becomes readable when its kernel-maintained counter is greater than zero. A write does not enqueue a variable-length message. It adds an unsigned 64-bit value to that counter, turning accumulated notification state into ordinary file-descriptor readiness. This boundary is useful in systems where a thread or kernel facility must wake an event loop without introducing a byte-stream protocol. The state carried by the descriptor is deliberately narrow: a counter, a readiness condition, and two possible consumption semantics.

Software Engineering 18 Sep 2026 6 min read

EPOLLEXCLUSIVE Limits Wakeups Across Competing epoll Instances

EPOLLEXCLUSIVE changes which epoll waiters are awakened when several epoll instances monitor the same target. Without the flag, a readiness event can be delivered to every attached epoll instance. With exclusive registration, Linux can wake a smaller subset, reducing redundant scheduling in configurations that otherwise create a thundering herd. The flag changes wakeup distribution. It does not assign permanent ownership of the target descriptor, serialize I/O, or guarantee that exactly one application thread consumes each unit of work.

Software Engineering 18 Sep 2026 4 min read

CLOSE_RANGE_UNSHARE Isolates Descriptor Table Cleanup

A thread preparing to cross an execve() boundary can need to remove every file descriptor above a small preserved set while other threads still share its descriptor table. Closing descriptors one by one creates a race: another thread can allocate a descriptor into the interval while cleanup is in progress. Linux close_range() with CLOSE_RANGE_UNSHARE changes the table-sharing boundary before applying the range operation. This behavior matters because a file descriptor number is only an index into a process descriptor table. With CLONE_FILES, multiple tasks can refer to the same table, so a close performed through one task changes descriptor visibility for all tasks sharing it. CLOSE_RANGE_UNSHARE gives the calling task a private descriptor table as part of the operation.

Software Engineering 17 Sep 2026 6 min read

userfaultfd Moves Missing-Page Resolution into User Space

A thread can touch a valid virtual address and stop before the access completes because the page has no present backing yet. With a range registered in UFFDIO_REGISTER_MODE_MISSING, Linux can report that fault through userfaultfd instead of resolving it entirely inside the kernel. A user-space manager then decides which page contents become visible before the blocked access resumes. This changes the ownership of one part of page-fault handling. The kernel still detects the fault, validates the virtual memory area, blocks the faulting execution, and installs mappings through the UFFDIO_* interface. User space gains control over the content and timing of resolution for registered faults.

Software Engineering 17 Sep 2026 4 min read

timerfd Counts Expirations Through Descriptor Reads

A periodic timerfd can expire several times before user space reads it. The next successful read() does not report only the most recent tick: it returns an unsigned 64-bit count of expirations accumulated since the previous successful read, or since the timer was configured if no read has completed yet. That behavior makes a timer an event-loop object without converting each expiration into a signal. The descriptor becomes readable when at least one expiration is pending, and the same descriptor can participate in poll(), select(), or epoll() beside sockets, pipes, and other descriptor-backed event sources.

Software Engineering 17 Sep 2026 6 min read

SO_REUSEPORT Moves TCP Connection Distribution Into the Kernel

With SO_REUSEPORT, several Linux TCP sockets can listen on the same local address and port at the same time. Incoming connections are assigned to a member of that reuseport group before an application calls accept(). The application no longer needs one shared listening socket as the sole handoff point between the network stack and multiple workers. That changes more than bind eligibility. It moves connection distribution into the kernel and gives each listener its own socket identity and accept path. The resulting architecture has different queueing, lifecycle, and routing properties from a design in which many workers compete on one listening socket.

Software Engineering 17 Sep 2026 4 min read

SO_REUSEPORT Forms Kernel-Selected Socket Groups

Multiple Linux sockets can bind the same local address when every participating socket enables SO_REUSEPORT before bind(). Incoming traffic is then assigned to a member of the resulting reuseport group rather than delivered to every socket. The shared address is therefore a kernel selection boundary, not a broadcast endpoint. This behavior supports independent receive or accept loops without forcing all work through one listening descriptor. It also creates a distinct operational property: group membership and the selection policy determine which socket receives a packet or connection.

Software Engineering 17 Sep 2026 5 min read

signalfd Turns Pending Signals into Readable Records

A signal included in a signalfd mask can make a file descriptor readable instead of invoking an asynchronous handler, provided that signal is blocked from ordinary delivery. A successful read() then consumes pending signal state and returns one or more signalfd_siginfo records. This changes the interface used to receive selected signals, but it does not replace Linux signal semantics. Signal masks, process-directed versus thread-directed delivery, standard-signal coalescing, and the special status of SIGKILL and SIGSTOP still define the boundary around the descriptor.

Software Engineering 17 Sep 2026 5 min read

signalfd Consumes Blocked Signals Through Descriptor Reads

A Linux signalfd becomes readable when a signal selected by its mask is pending for the reading context. A successful read() does more than observe that state: it consumes the returned signal occurrences, removing them from pending signal state. That behavior gives signals a descriptor-facing consumption path. It does not convert the signal subsystem into a byte stream, and it does not replace the signal mask that controls ordinary delivery.

Software Engineering 17 Sep 2026 6 min read

Seccomp User Notification Delegates Selected System Calls to a Supervisor

A system call selected by a seccomp filter can stop before kernel execution and appear as a request on a notification file descriptor. With SECCOMP_RET_USER_NOTIF, Linux turns that call into a coordination point between the blocked target and a userspace supervisor. The supervisor can emulate a result, inject a file descriptor for suitable operations, or permit the kernel to continue the original call. This mechanism is deliberately narrower than a general userspace security policy engine. Its strongest boundary is the kernel-mediated suspension and response protocol. Data reached through target-memory pointers can still change around a supervisor’s inspection, and a response that continues the original call re-enters ordinary kernel execution with that race still relevant.

Software Engineering 17 Sep 2026 4 min read

renameat2 RENAME_EXCHANGE Swaps Two Paths in One Filesystem Operation

renameat2() with RENAME_EXCHANGE changes two existing directory entries as one atomic rename operation. Before the call, each pathname reaches its original object; after a successful call, each pathname reaches the object formerly named by the other path. There is no successful intermediate state in which one of the two names has merely been removed or overwritten. That property is distinct from ordinary rename(). A conventional rename can atomically replace a destination, but replacement discards the destination name from the namespace. Exchange preserves both named objects and swaps their positions.

Software Engineering 17 Sep 2026 4 min read

pidfd Keeps Process Identity Stable Across PID Reuse

A numeric Linux PID can be reused after its process exits. A PID file descriptor instead refers to a specific task, so later operations through that descriptor do not silently retarget a different process that receives the same numeric PID. This changes process identity from a lookup repeated at each operation into a kernel-held reference with descriptor semantics. The distinction matters for signaling, exit monitoring, and event loops that retain process handles across asynchronous work.