Skip to content

Archive

File Descriptors

35 articles
Linux 23 Sep 2026 4 min read

TFD_TIMER_CANCEL_ON_SET Turns Realtime Clock Jumps into ECANCELED

An absolute timer tied to CLOCK_REALTIME has a dependency that a monotonic deadline does not: an administrator, synchronization service, or privileged process can move the wall clock discontinuously while the timer is armed. Linux timerfd can expose that event explicitly. With TFD_TIMER_ABSTIME | TFD_TIMER_CANCEL_ON_SET, a qualifying clock jump causes a current or later read() on the timer descriptor to fail with ECANCELED. This behavior separates two events that would otherwise be easy to conflate: reaching a scheduled wall-clock instant and invalidating the clock basis used to schedule it.

Linux 23 Sep 2026 4 min read

pidfd_getfd Duplicates a Target Descriptor into the Calling Process

pidfd_getfd Duplicates a Target Descriptor into the Calling Process pidfd_getfd() can place a duplicate of another process’s open file descriptor into the caller’s descriptor table. The returned descriptor is local to the caller, but it refers to the same open file description as the selected descriptor in the target process. That distinction matters. This operation does not reopen a pathname, reconstruct a socket, or create an independent file position. It duplicates an existing kernel reference across a process boundary.

Linux 23 Sep 2026 5 min read

eventfd Semaphore Mode Turns Counter Values into Single-Unit Reads

Linux eventfd exposes a kernel-maintained 64-bit counter through a file descriptor. Its compact interface hides an important semantic choice: a normal read drains the current counter value, while an EFD_SEMAPHORE read consumes exactly one unit. The write path still adds values to the same counter. That distinction changes the object from an aggregate notification counter into a descriptor-backed source of individually consumable units. The readiness model remains compatible with poll, epoll, and related descriptor multiplexing, so the same object can connect producer accounting with an event loop without adding a separate pipe payload.

Linux 23 Sep 2026 5 min read

close_range with UNSHARE Separates File Descriptor Cleanup from Peer Threads

A multithreaded Linux process can share one file descriptor table across its threads. That arrangement is convenient during normal execution: a descriptor opened by one thread becomes available to its peers. It becomes less convenient when one thread is preparing a restricted execution context and wants to discard a broad descriptor range without racing with peers that can still allocate descriptors. close_range() provides a range operation for this boundary. With CLOSE_RANGE_UNSHARE, the kernel first separates the caller from the shared descriptor table and then applies the requested closure to the caller’s resulting table. The operation is close to combining unshare(CLONE_FILES) with a range close, but the kernel can perform less work in common cases.

Linux 20 Sep 2026 5 min read

pidfd_getfd Duplicates a Live File Descriptor Across Processes

A process can acquire a usable duplicate of a file descriptor that is already open in another process without asking that process to send it over a UNIX domain socket. Linux pidfd_getfd() performs that transfer through a PID file descriptor, subject to a ptrace access check. The returned descriptor is new in the caller, but the kernel object behind it is not independent. It refers to the same open file description as the target descriptor. That distinction controls offset sharing, file status flags, and operations on the underlying object.

Linux 19 Sep 2026 5 min read

timerfd Reads Count Periodic Expirations

A periodic timer can expire several times before a busy event loop gets CPU time again. Linux timerfd does not compress that delay into a bare “timer fired” notification. A successful read() returns an unsigned 64-bit count of expirations accumulated since the timer was armed or since the preceding successful read. That counter changes the semantics of delayed timer handling. Readiness says at least one expiration is pending; the value read from the descriptor says how many periods elapsed.

Linux 19 Sep 2026 6 min read

pidfd_getfd Duplicates a Live Descriptor Across Process Boundaries

A process can acquire a new descriptor that refers to the same open file description as a descriptor already held by another process, without asking that target process to send it. Linux provides this operation through pidfd_getfd(). The resulting descriptor is local to the caller, but the kernel object behind it is shared with the target descriptor. That distinction matters because a descriptor number is only an entry in one process’s descriptor table. The open file description carries state such as the current file offset and file status flags. Duplicating across a process boundary therefore transfers access to an existing kernel file instance rather than reopening the pathname or constructing an independent instance.

Linux 19 Sep 2026 4 min read

memfd Seals Constrain Shared-Memory Mutation After Handoff

A memfd_create() descriptor names an anonymous file whose storage lives in memory-backed filesystem infrastructure. By itself, descriptor handoff does not freeze that object: a process retaining suitable access can still write bytes, truncate the file, or extend it. Linux file seals add kernel-enforced restrictions that can make selected mutations fail after the producer declares the object complete. This changes shared-memory handoff from a convention into a state transition enforced at the file object.

Linux 19 Sep 2026 5 min read

eventfd Turns Kernel Notifications into Pollable Counters

A Linux process can signal work through a file descriptor without moving a byte stream between producer and consumer. eventfd() creates a kernel-maintained 64-bit counter whose readiness can be observed by poll(), select(), or epoll. A write adds to the counter; a read consumes its accumulated state according to the descriptor mode. That shape makes eventfd different from a pipe. A pipe preserves a sequence of bytes. An eventfd preserves counter state. When the application needs a wakeup edge plus a compact amount of accumulated state, that distinction removes buffering and framing that a byte stream would otherwise require.

Linux 19 Sep 2026 5 min read

close_range Makes File-Descriptor Cleanup a Single Linux Operation

A process preparing to execute another program often needs a simple boundary: descriptors 0, 1, and 2 remain available, while every higher descriptor must disappear. Repeating close() over a guessed numeric limit or enumerating /proc/self/fd turns that boundary into a userspace scan. Linux close_range() expresses the interval directly. The kernel applies one operation to every open file descriptor from first through last, inclusive. With flags, the same interface can isolate a shared descriptor table or mark the interval close-on-exec instead of closing it immediately.

Software Engineering 18 Sep 2026 6 min read

SCM_RIGHTS Transfers Open File Descriptions Across Process Boundaries

SCM_RIGHTS lets one process send a reference to an open file through a Unix domain socket. The receiver obtains a file descriptor in its own descriptor table, but the transfer does not reopen the pathname or copy the kernel object. On Linux, the resulting reference has semantics equivalent to duplicating the sender’s descriptor into the receiving process. That distinction matters whenever a process boundary is also an authority boundary. A supervisor can open a socket, file, pipe, device, or other descriptor-backed object and pass the established reference to a worker. The worker receives access to the already-open object, including open-file state that can remain shared with the sender.

Linux 18 Sep 2026 6 min read

pidfd_getfd Duplicates Another Process File Descriptor into the Caller

A file descriptor number has meaning only inside its process descriptor table, but the kernel object behind that number can be shared across processes. Linux pidfd_getfd() bridges those two scopes: it takes a PID file descriptor plus a descriptor number from the referenced process and installs a duplicate descriptor in the caller. The new descriptor refers to the same open file description as the target descriptor. That last property is the central boundary. pidfd_getfd() does not reopen a pathname, copy bytes, or create an independent file position. It duplicates an existing kernel reference and therefore inherits sharing semantics that can affect both processes.

Cybersecurity 18 Sep 2026 4 min read

Memfd Seals Turn Mutable Anonymous Files into Explicit Handoff Objects

A process prepares a binary payload in memory, passes a file descriptor to another process, and expects the bytes to remain stable after validation. A plain descriptor does not create that guarantee. If some holder still has write authority, the object can change after a consumer has inspected it, and pathname permissions offer no useful boundary when the object has no ordinary filesystem name. Linux memfd_create() provides an anonymous file backed by memory-like filesystem storage, and file seals can constrain later changes to that file. The useful security property is not anonymity by itself. It is the ability to construct a mutable object, apply irreversible restrictions to that object, then hand out descriptors whose backing file can no longer be changed in the prohibited ways.

Cybersecurity 18 Sep 2026 6 min read

memfd File Seals Turn Shared Memory into a Kernel-Enforced Mutation Boundary

A broker can allocate a memory-backed object, populate it, and pass its file descriptor to another process over a UNIX domain socket. The receiver may treat the bytes as immutable configuration, compiled code, or a serialized artifact. That assumption is unsafe if the sender or another holder can still alter the same inode after validation. Linux memfd_create() and file seals provide a kernel-enforced way to narrow that mutation surface without assigning the object a persistent filesystem pathname.

Software Engineering 18 Sep 2026 6 min read

Linux pidfd Binds Process Operations to Stable Kernel References

A numeric process ID is a name in a PID namespace, not a durable handle to one process lifetime. After a process exits and its PID becomes available for reuse, a later process can receive the same number. Linux PID file descriptors add a different interface boundary: a pidfd is a file descriptor referring to a particular task, so later operations can target that kernel reference rather than resolving the numeric PID again.

Software Engineering 18 Sep 2026 5 min read

Linux O_PATH Separates Object Reference From I/O Authority

open() usually combines two effects: pathname resolution selects a filesystem object, then the returned file descriptor carries an access mode for data I/O. Linux O_PATH splits those effects. A successful open(path, O_PATH) returns a descriptor that refers to the selected object while ordinary read() and write() through that descriptor are not permitted. That split is useful anywhere a process needs a durable kernel reference for later metadata or pathname-relative operations without opening the object for data transfer. It also changes race analysis: later operations can start from the descriptor rather than resolving the original pathname again.

Software Engineering 18 Sep 2026 4 min read

Linux memfd Seals Turn Mutable Memory Files into Enforced State Transitions

A file created by memfd_create() can begin as mutable storage and later acquire kernel-enforced restrictions that apply to the underlying file rather than to one descriptor. With MFD_ALLOW_SEALING, a process can add seals through fcntl(F_ADD_SEALS) and make selected mutations unavailable to every holder of that file. This creates a state transition that ordinary descriptor permissions do not express. A producer can populate bytes, fix the file’s size, and then publish the descriptor with restrictions that remain attached even after the descriptor crosses a process boundary.

Software Engineering 18 Sep 2026 6 min read

Linux memfd Seals Convert Mutable Anonymous Files into Restricted Capabilities

A file descriptor returned by memfd_create() can begin as a writable, resizable anonymous file and later become an object whose permitted mutation operations have been permanently reduced. Linux implements that transition with file seals. The mechanism is attached to the underlying file rather than to one descriptor, so passing a duplicate descriptor across a process boundary does not create an independent sealing state. This property makes sealing more than a convenience around temporary storage. It changes the authority carried by every descriptor that refers to the same memfd object. The transition is monotonic: seals can be added, but they cannot be removed.

Software Engineering 18 Sep 2026 5 min read

Linux inotify Reports Directory-Entry Events, Not Durable Path Identity

An inotify watch does not make a pathname a durable identifier. Linux attaches a watch to a filesystem object selected when inotify_add_watch() succeeds, then emits records describing activity associated with watched objects and directory entries. Names can move, objects can disappear, and event delivery can lose detail when the queue overflows. That boundary matters for file synchronizers, configuration reloaders, indexers, and service supervisors. An event stream can signal that local filesystem state changed, but reconstructing authoritative state still depends on filesystem operations performed after the event.

Software Engineering 18 Sep 2026 8 min read

Linux close_range Makes Descriptor-Table Cleanup a Range Operation

A process preparing to execute another program often needs a simple descriptor invariant: standard input, output, and error remain available, while unrelated descriptors do not cross the execution boundary. Closing descriptors one at a time can turn that invariant into an enumeration problem. Linux close_range() instead applies an operation to an inclusive numeric interval in the calling task’s file-descriptor table. The interface is small, but its semantics reach into descriptor-table sharing, execve() inheritance, concurrent descriptor allocation, and privilege transitions. The flags select more than implementation strategy: they determine whether descriptors disappear immediately, become close-on-exec, or are first separated from a table shared with other tasks.

Software Engineering 18 Sep 2026 5 min read

Linux close_range Makes Descriptor Cleanup a Table Operation

A process preparing for execve() often needs a simple invariant: descriptors above a small allowlist must not survive into the new program. Closing descriptor numbers one by one turns that invariant into an enumeration problem. Linux close_range() expresses it directly as an operation over an inclusive interval of the calling task’s file descriptor table. The interface is Linux-specific. Its behavior belongs to Linux file-table and system-call semantics, not to the C language or a portable POSIX guarantee.

Software Engineering 18 Sep 2026 4 min read

CLOSE_RANGE_UNSHARE Isolates Descriptor Table Cleanup

A thread preparing to cross an execve() boundary can need to remove every file descriptor above a small preserved set while other threads still share its descriptor table. Closing descriptors one by one creates a race: another thread can allocate a descriptor into the interval while cleanup is in progress. Linux close_range() with CLOSE_RANGE_UNSHARE changes the table-sharing boundary before applying the range operation. This behavior matters because a file descriptor number is only an index into a process descriptor table. With CLONE_FILES, multiple tasks can refer to the same table, so a close performed through one task changes descriptor visibility for all tasks sharing it. CLOSE_RANGE_UNSHARE gives the calling task a private descriptor table as part of the operation.

Cybersecurity 18 Sep 2026 6 min read

close_range Narrows File Descriptor Inheritance Before exec

A service process can accumulate sockets, pipes, directory handles, log files, and control descriptors long before it launches a helper. If those descriptors survive into the new program, the helper receives capabilities that its command-line arguments and environment do not reveal. A connected socket can carry authenticated access; an open directory can preserve reachability to a filesystem location; a pipe can expose another component’s data path. Linux close_range() gives pre-exec code a range operation over file descriptors. Its security value is not that descriptors become harmless. It is that a process can narrow the descriptor set that crosses an execve() boundary without enumerating /proc/self/fd or issuing one close() call per candidate descriptor.

Linux 18 Sep 2026 6 min read

close_range Controls Descriptor Inheritance Across exec Boundaries

A process preparing to call execve() may have hundreds or thousands of open file descriptors, while the new program should inherit only a small selected set. Closing descriptors one by one creates both bookkeeping cost and a concurrency problem: another thread can allocate a descriptor while the cleanup loop is still running. Linux close_range() moves that operation to the descriptor-table boundary. A caller specifies an inclusive numeric range and asks the kernel either to close descriptors in that range or mark them close-on-exec. With CLOSE_RANGE_UNSHARE, the caller can first detach its descriptor table from threads or processes that share it.