Skip to content

Topic archive

Linux

An open-source operating system ecosystem known for its flexibility, developer tooling, and widespread use on servers, cloud platforms, embedded devices, and workstations.

115 articles
Linux 18 Sep 2026 6 min read

pidfd_getfd Duplicates Another Process File Descriptor into the Caller

A file descriptor number has meaning only inside its process descriptor table, but the kernel object behind that number can be shared across processes. Linux pidfd_getfd() bridges those two scopes: it takes a PID file descriptor plus a descriptor number from the referenced process and installs a duplicate descriptor in the caller. The new descriptor refers to the same open file description as the target descriptor. That last property is the central boundary. pidfd_getfd() does not reopen a pathname, copy bytes, or create an independent file position. It duplicates an existing kernel reference and therefore inherits sharing semantics that can affect both processes.

Linux 18 Sep 2026 5 min read

openat2 Resolve Flags Constrain Path Traversal per Open

A pathname passed to openat2() can be rejected even when the same pathname would resolve successfully through openat(). The difference comes from open_how.resolve: Linux can apply traversal constraints while resolving every component of that single open operation. This changes the boundary around path handling. A directory file descriptor can act as more than a starting point; resolve flags can restrict escapes, symbolic-link traversal, mount crossings, and lookups that require work beyond cached state.

Linux 18 Sep 2026 5 min read

mseal Locks Memory Mapping Layout and Permissions

A process can establish a memory mapping with the intended address, size, and protection bits, then later alter that mapping with operations such as munmap(), mprotect(), or mremap(). Linux mseal() adds a one-way state transition: selected virtual memory areas can be sealed so a class of later mapping modifications is rejected by the kernel. The mechanism protects mapping structure rather than the bytes stored in the mapping. A writable sealed mapping remains writable through ordinary stores. Sealing instead constrains operations that could remove the mapping, relocate it, replace it, or change attributes covered by the sealing rules.

Linux 18 Sep 2026 6 min read

membarrier Moves Memory-Ordering Cost to an Infrequent Coordination Path

A full hardware memory barrier in a frequently executed path can impose a cost on every operation, even when cross-thread coordination happens only occasionally. Linux membarrier() supports a different placement of that cost: a rare coordination path can request an ordering event across a defined set of threads while a frequent path may need only compiler-level ordering. This is not a general replacement for atomics, locks, or the memory model of a programming language. It is a Linux-specific synchronization primitive for designs whose correctness already has a precise pairing between a frequent path and an infrequent coordination path.

Linux 18 Sep 2026 6 min read

Landlock Rulesets Add Process-Local Access Control

A Linux process can voluntarily remove access that its UID, capabilities, mount namespace, and other security layers would otherwise permit. Landlock implements this as a stackable Linux Security Module: a process creates a ruleset, adds allowed objects, then places itself in a Landlock domain. The resulting policy is an additional restriction. It does not grant access denied by DAC, ACLs, SELinux, AppArmor, mount permissions, or another active control. Once enforced, the Landlock layer cannot be removed from that thread; later Landlock domains can only add restrictions.

Linux 18 Sep 2026 5 min read

eventfd Aggregates Notifications in a Kernel Counter

An eventfd can absorb several notification writes before userspace services the descriptor. The kernel stores those writes in a 64-bit counter, so readiness represents pending counter state rather than a queue containing one record per notification. That distinction matters in event loops. A producer can add values while a consumer is occupied, and the next read can collapse accumulated state into one result. With EFD_SEMAPHORE, the same object exposes a different consumption rule without changing its readiness model.

Linux 18 Sep 2026 4 min read

EPOLLEXCLUSIVE Limits Wakeups Across Competing epoll Instances

A single ready socket can wake several threads when each thread waits on a different epoll instance that watches that socket. Linux provides EPOLLEXCLUSIVE to narrow that wakeup fan-out: among epoll instances that registered the target with the flag, a readiness event wakes one or more rather than all of them. The distinction is deliberately weaker than “exactly one waiter.” EPOLLEXCLUSIVE changes notification selection across epoll instances. It does not transfer ownership of the file descriptor, serialize all I/O, or guarantee that only one thread can observe useful work.

Linux 18 Sep 2026 6 min read

close_range Controls Descriptor Inheritance Across exec Boundaries

A process preparing to call execve() may have hundreds or thousands of open file descriptors, while the new program should inherit only a small selected set. Closing descriptors one by one creates both bookkeeping cost and a concurrency problem: another thread can allocate a descriptor while the cleanup loop is still running. Linux close_range() moves that operation to the descriptor-table boundary. A caller specifies an inclusive numeric range and asks the kernel either to close descriptors in that range or mark them close-on-exec. With CLOSE_RANGE_UNSHARE, the caller can first detach its descriptor table from threads or processes that share it.

Linux 17 Sep 2026 4 min read

Truncating a Mapped File Can Trigger SIGBUS

A process can retain a valid mmap() address range after another operation shrinks the backing file, then receive SIGBUS when it touches a mapped page past the file’s new end. The mapping itself has not vanished. Its backing object no longer covers every page that the virtual mapping originally referenced. This boundary is easy to miss because mapping lifetime and file size are separate state. Closing the original file descriptor does not invalidate an established mapping, and shrinking the file does not act like munmap() on every process that maps it.

Linux 17 Sep 2026 5 min read

TCP_NODELAY Disables Nagle Coalescing on a Socket

A TCP socket can hold a small write instead of transmitting it immediately when earlier data remains unacknowledged. This behavior comes from Nagle coalescing: it limits the stream of small TCP segments by allowing outstanding data to influence transmission of newly queued bytes. On Linux, setting TCP_NODELAY disables that coalescing rule for the socket. Small writes become eligible for prompt transmission, subject to the rest of the TCP stack, congestion control, flow control, queue state, and device scheduling.

Linux 17 Sep 2026 6 min read

SO_REUSEPORT Distributes Traffic Across Socket Groups

SO_REUSEPORT changes a local endpoint from a single-socket binding into a socket group. On Linux, multiple TCP or UDP sockets can bind the same local address when every participating socket enables the option before bind() and the bind credentials satisfy the kernel’s reuse rules. That behavior is distinct from merely relaxing address-conflict checks. Incoming traffic must also be assigned to one member of the group. The resulting selection boundary affects listener architecture, queue isolation, process restarts, UDP flow placement, and any design that assumes a port maps to exactly one socket.

Linux 17 Sep 2026 4 min read

SO_RCVLOWAT Raises the Readability Threshold for Linux Sockets

A Linux socket with SO_RCVLOWAT set above one byte can have data queued while poll(), select(), or epoll still reports no normal readable readiness. Since Linux 2.6.28, those readiness interfaces respect the configured receive low-water mark. The option changes the threshold associated with normal receive readiness. It does not define message boundaries, reserve receive-buffer space, or guarantee that a later receive operation returns exactly the configured number of bytes. Readability can require more than one queued byte Socket receive readiness is usually observed with the default low-water mark of one byte. In that state, ordinary queued data is enough to satisfy the data-volume part of the readable condition.

Linux 17 Sep 2026 5 min read

rename Replaces a Directory Entry Atomically on Linux

A successful rename() can replace an existing destination pathname without exposing an intermediate state in which that destination name is missing. Processes resolving the destination observe either the old directory entry or the replacement, subject to filesystem and mount constraints. That atomic namespace transition is narrower than several properties often associated with file replacement. It does not make prior writes durable, does not force directory metadata to stable storage, and does not invalidate file descriptors that already refer to the replaced file.

Linux 17 Sep 2026 4 min read

O_CLOEXEC Closes Descriptors Atomically Across exec

A file descriptor created without close-on-exec state can escape into a newly executed program during a narrow concurrency window. In a multithreaded Linux process, setting FD_CLOEXEC with a later fcntl() call leaves that window open between descriptor creation and the flag update. O_CLOEXEC removes the split operation. The kernel creates the descriptor with its close-on-exec flag already set, so another thread cannot observe an intermediate state in which the descriptor exists but remains inheritable across a successful execve().

Linux 17 Sep 2026 5 min read

O_APPEND Couples End Positioning with Each Write

O_APPEND changes a write from two separable actions into one coupled operation: Linux positions the open file description at the current end of the file and performs the write as a single atomic step. That property matters when multiple writers target one regular file. A sequence built from lseek(fd, 0, SEEK_END) followed by write(fd, ...) does not carry the same append semantics because another writer can change the file between those two system calls.

Linux 17 Sep 2026 4 min read

Duplicated File Descriptors Share an Open File Description

Two file descriptor numbers can move the same file offset. On Linux, this occurs when both descriptors refer to one open file description, as happens after dup() and across inherited descriptors after fork(). The distinction matters because a file descriptor is a process-visible integer, while the open file description is the kernel object that carries state for an open instance of a file. Treating those layers as interchangeable can produce offset interference, status-flag changes that cross descriptor boundaries, and surprising behavior after process creation.

Linux 16 Sep 2026 5 min read

Linux TCP TIME_WAIT Retains Closed Connection State

A TCP socket can disappear from an application while the kernel still retains state for the closed connection. On Linux, the endpoint that completes the active close commonly enters TIME_WAIT, keeping enough protocol state to protect a later connection from delayed segments associated with the old one. This state is not evidence that a process forgot to close a file descriptor. The application-visible socket can already be gone. TIME_WAIT belongs to TCP’s connection-lifecycle machinery and persists independently of the process that initiated the close.

Linux 16 Sep 2026 4 min read

Linux TCP Autocorking Coalesces Consecutive Small Writes

A small TCP write does not always trigger an immediate packet transmission on Linux. With TCP autocorking enabled, the stack can defer a new small send when an earlier packet from the same flow is still waiting in a qdisc or device transmit queue, giving a following write a chance to join the pending data. The mechanism targets packet count rather than application-visible buffering semantics. A successful write() or sendmsg() still reports bytes accepted by the socket; autocorking influences when queued bytes advance into transmission.

Linux 16 Sep 2026 6 min read

Linux Readahead Expands Sequential Page-Cache Reads

A buffered file read can cause Linux to fetch more data than the application explicitly requested. The extra I/O is readahead: the kernel populates nearby page-cache folios in anticipation of continued access. This behavior sits between application read size and storage request size. A process may issue modest read() calls while the kernel submits larger reads to keep later accesses from waiting on storage. Readahead is page-cache speculation Buffered file I/O normally passes through the page cache. When requested file data is absent, the kernel must arrange I/O for that miss. The readahead path can extend that operation across additional folios that are not yet present in the cache.

Linux 16 Sep 2026 5 min read

Linux PSI Separates Partial and Total Resource Stalls

A Linux host can report modest CPU utilization while runnable work is delayed, or ample memory capacity while tasks repeatedly stall in reclaim. Utilization counters describe resource activity; pressure stall information records time in which work cannot make progress because a resource is contended. PSI exposes that lost execution opportunity through CPU, memory, and I/O pressure files. Its central distinction is between a stall affecting at least one task and a stall that leaves every non-idle task unable to make progress.

Linux 16 Sep 2026 6 min read

Linux cgroup memory.high Converts Overage into Reclaim Pressure

A cgroup can remain alive after its memory usage crosses memory.high. The boundary does not behave like a hard allocation ceiling: tasks in the cgroup are throttled and pushed into heavy reclaim pressure, and usage can remain above the configured value under extreme conditions. That behavior makes memory.high materially different from memory.max. The former converts excess usage into execution cost and reclaim work. The latter is a hard limit that can lead to a cgroup OOM when reclaim cannot reduce usage enough.

Linux 16 Sep 2026 5 min read

io_uring Registered Files Bypass Repeated Descriptor Lookup

An io_uring request that uses a normal file descriptor still has to resolve that descriptor through the submitting task’s file table. A registered file takes a different path: the ring holds a reference to the open file, and an SQE names a slot in that ring-local table. That distinction removes repeated descriptor lookup from the request path. It also changes resource lifetime, update semantics, and the meaning of the SQE fd field.

Linux 16 Sep 2026 5 min read

epoll Edge-Triggered Readiness Requires Draining

An edge-triggered epoll consumer can block while unread data is still buffered. The failure appears when an event is consumed, only part of the available input is read, and the event loop returns to epoll_wait() expecting another notification for the bytes that remain. This behavior follows directly from EPOLLET. Edge-triggered notification reports changes in readiness rather than continuously reporting a ready condition. Once a readiness transition has produced an event, leaving the file descriptor ready does not itself create a new transition.

Linux 07 Sep 2026 12 min read

Understand Sparse Files on Linux Without Wasting Disk Space

A file can report a size of 100 GiB without consuming 100 GiB of disk blocks. That sounds contradictory until you separate two ideas that ordinary file APIs often present together: a file’s logical size and the storage that the filesystem has actually allocated for it. Linux filesystems can represent long ranges of unwritten bytes as holes. Reading a hole returns zero bytes, but the filesystem does not need to store one physical zero byte for every logical byte in that range. A regular file that contains holes is called a sparse file.