Skip to content

Archive

Filesystems

29 articles
Tech 22 Sep 2026 5 min read

Linux Buffered Writes Separate Write Completion from Storage Persistence

A successful buffered write() does not generally mean that the new file data has already reached non-volatile storage. On Linux, the common buffered I/O path places file data in the page cache, marks the affected cache state dirty, and lets storage I/O occur later. That separation is central to normal filesystem I/O. Memory absorbs application writes at CPU-accessible speed, while the kernel can schedule backing-device traffic independently. The result improves flexibility and can reduce immediate storage stalls, but it also creates a boundary between syscall completion and persistence.

Linux 21 Sep 2026 5 min read

openat2 Resolve Flags Constrain Path Traversal During Lookup

A pathname passed to openat() is resolved by the kernel, but the caller has limited control over traversal through intermediate components. Linux openat2() adds a resolve field that applies constraints to the complete lookup operation. The restriction is evaluated while components are traversed, rather than by validating a pathname in userspace and opening it later. That distinction matters when path components can change concurrently. A userspace sequence that checks a path and then opens it creates separate observations of mutable filesystem state. openat2() places the selected lookup policy in the same system call that returns the file descriptor.

Linux 19 Sep 2026 5 min read

openat2 Constrains Linux Path Resolution at the Open Boundary

A pathname can change meaning while a process is resolving it. Directory renames, symbolic links, mount points, and .. components can redirect lookup away from the directory a program intended to treat as its boundary. Linux openat2() attaches resolution policy to the lookup itself. Its struct open_how contains a resolve bit mask, so the kernel can reject a path when resolution violates a caller-selected constraint instead of relying only on checks performed before open().

Software Engineering 19 Sep 2026 8 min read

Open File Description Locks Bind Byte Ranges to File Instances

A byte-range lock can protect the same inode yet have radically different lifetime semantics depending on what owns the lock. Traditional fcntl() record locks are process-associated. Open file description locks instead attach to the kernel open file description referenced by a descriptor. That shift changes which close operation releases a lock, what survives fork(), and whether two threads in one process can contend on the same file region. Linux exposes this model through F_OFD_SETLK, F_OFD_SETLKW, and F_OFD_GETLK. The range model remains familiar: struct flock specifies a read lock, write lock, or unlock together with an offset and length. The significant difference is ownership.

Linux 19 Sep 2026 6 min read

Idmapped Mounts Remap Ownership Without Rewriting Inodes

The same inode can appear with different ownership through two mount points without any recursive chown(). Linux idmapped mounts attach an ID mapping to a mount, so VFS ownership presentation and permission checks can translate user and group IDs for that view while the ownership stored by the filesystem remains unchanged. This property separates persistent inode metadata from the identity view exposed at a particular mount. It is especially useful when a filesystem tree must be shared with a container whose user namespace maps IDs differently from the host.

Cybersecurity 19 Sep 2026 7 min read

fanotify Permission Events Put File Access Behind a User-Space Decision

A process can pass ordinary filesystem permission checks and still wait before its file operation completes. Linux fanotify permission events let a monitoring group intercept selected operations and require a user-space listener to return an allow or deny response. The mechanism inserts a synchronous decision point into the access path rather than merely reporting activity after it occurs. That distinction makes fanotify useful for security products that need content inspection or policy evaluation close to file access. It also creates a dependency that ordinary notification systems do not have: the kernel may be holding another process at a permission event while user space decides its fate.

Linux 18 Sep 2026 5 min read

openat2 Resolve Flags Constrain Path Traversal per Open

A pathname passed to openat2() can be rejected even when the same pathname would resolve successfully through openat(). The difference comes from open_how.resolve: Linux can apply traversal constraints while resolving every component of that single open operation. This changes the boundary around path handling. A directory file descriptor can act as more than a starting point; resolve flags can restrict escapes, symbolic-link traversal, mount crossings, and lookups that require work beyond cached state.

Software Engineering 18 Sep 2026 8 min read

MAP_SHARED mmap Couples Memory Writes to File-Backed Page State

A writable MAP_SHARED mapping lets a process modify file-backed state with ordinary memory stores. The bytes are addressed through virtual memory rather than passed to write(), but the mapping still participates in filesystem state: modifications can become visible through other shared mappings and file I/O, and dirty pages can later be written back to storage. That interface compresses several mechanisms into one address range. CPU stores, page faults, page-cache residency, filesystem writeback, and storage persistence can all participate in the lifetime of the same bytes. Treating a successful store as equivalent to durable file output collapses boundaries that the operating system keeps distinct.

Software Engineering 18 Sep 2026 6 min read

Linux renameat2 Makes Path-Replacement Policy Atomic

A pathname rename changes directory entries while open file descriptors continue to refer to the same underlying objects. Linux renameat2() adds policy to that namespace update: a caller can reject replacement, exchange two existing names, or request a whiteout for union-filesystem operation. These policies are executed as part of the rename operation rather than as checks performed separately in userspace. The interface is Linux-specific. A zero flags argument gives renameat() behavior, while nonzero flags add Linux semantics that also depend on support from the mounted filesystem.

Software Engineering 18 Sep 2026 6 min read

Linux openat2 Constrains Path Resolution Inside a Directory Boundary

A pathname passed to openat2() can be resolved relative to a directory file descriptor while the kernel enforces restrictions on the resolution process itself. That distinction matters when a process accepts path components from a less-trusted source. A string check can inspect the pathname text, but it cannot by itself freeze the filesystem namespace while lookup proceeds. openat2() places the policy beside the lookup. Its struct open_how separates ordinary open flags from resolve flags that constrain traversal. The resulting boundary is about resolution semantics, not merely the spelling of a path.

Software Engineering 18 Sep 2026 5 min read

Linux inotify Reports Directory-Entry Events, Not Durable Path Identity

An inotify watch does not make a pathname a durable identifier. Linux attaches a watch to a filesystem object selected when inotify_add_watch() succeeds, then emits records describing activity associated with watched objects and directory entries. Names can move, objects can disappear, and event delivery can lose detail when the queue overflows. That boundary matters for file synchronizers, configuration reloaders, indexers, and service supervisors. An event stream can signal that local filesystem state changed, but reconstructing authoritative state still depends on filesystem operations performed after the event.

Software Engineering 18 Sep 2026 8 min read

Linux copy_file_range Separates Copy Semantics From Data Movement

copy_file_range() asks Linux to copy bytes between regular files without requiring the application to shuttle those bytes through a user-space buffer. The call defines a byte-range operation, but it does not prescribe the physical transfer mechanism. A filesystem can perform ordinary data movement, use a copy-on-write sharing mechanism such as reflink, or employ another supported acceleration path while preserving the visible file contents required by the operation. That separation is the central API boundary. Applications specify source and destination ranges and observe the returned byte count. The kernel and filesystem retain latitude over the mechanism used to realize the copy.

Linux 17 Sep 2026 4 min read

Truncating a Mapped File Can Trigger SIGBUS

A process can retain a valid mmap() address range after another operation shrinks the backing file, then receive SIGBUS when it touches a mapped page past the file’s new end. The mapping itself has not vanished. Its backing object no longer covers every page that the virtual mapping originally referenced. This boundary is easy to miss because mapping lifetime and file size are separate state. Closing the original file descriptor does not invalidate an established mapping, and shrinking the file does not act like munmap() on every process that maps it.

Software Engineering 17 Sep 2026 4 min read

renameat2 RENAME_EXCHANGE Swaps Two Paths in One Filesystem Operation

renameat2() with RENAME_EXCHANGE changes two existing directory entries as one atomic rename operation. Before the call, each pathname reaches its original object; after a successful call, each pathname reaches the object formerly named by the other path. There is no successful intermediate state in which one of the two names has merely been removed or overwritten. That property is distinct from ordinary rename(). A conventional rename can atomically replace a destination, but replacement discards the destination name from the namespace. Exchange preserves both named objects and swaps their positions.

Linux 17 Sep 2026 5 min read

rename Replaces a Directory Entry Atomically on Linux

A successful rename() can replace an existing destination pathname without exposing an intermediate state in which that destination name is missing. Processes resolving the destination observe either the old directory entry or the replacement, subject to filesystem and mount constraints. That atomic namespace transition is narrower than several properties often associated with file replacement. It does not make prior writes durable, does not force directory metadata to stable storage, and does not invalidate file descriptors that already refer to the replaced file.

Software Engineering 17 Sep 2026 5 min read

openat2 Constrains Path Resolution at the Lookup Boundary

A pathname can begin below a trusted directory and still escape that subtree during resolution. A .. component, symbolic link, magic link, or mount transition can change the object ultimately reached even when the initial directory file descriptor is trusted. Linux openat2() places constraints inside pathname resolution itself, so the kernel can reject a lookup that violates the selected boundary. This differs from checking a pathname string before calling open(). Path resolution operates on filesystem objects and namespace state, not only text. openat2() extends the openat() model with a struct open_how whose resolve field controls traversal of pathname components.

Software Engineering 17 Sep 2026 6 min read

Linux openat2 Makes Path Resolution Policy Part of the Open Operation

A pathname is not an object reference. It is an instruction for traversing a mutable namespace, and another task can alter directory entries, symbolic links, or mounts while that traversal is relevant to an application. Linux openat2() addresses this boundary by placing path-resolution constraints in the same kernel operation that returns the file descriptor. That placement matters when a program accepts a pathname but intends to confine resolution to a directory tree. A user-space sequence that inspects components and later calls open() separates validation from use. openat2() can instead make selected traversal rules part of the lookup itself.

Software Engineering 17 Sep 2026 9 min read

Linux openat2 Makes Path Resolution Constraints Atomic

A pathname can name a different object by the time a second lookup checks it. On Linux, openat2() addresses that boundary by attaching resolution constraints to the same kernel operation that walks the pathname and opens the resulting object. The policy is evaluated during lookup rather than inferred from a pathname inspected before or after the open. This distinction matters whenever a process accepts path components from a less-trusted source while intending to keep resolution inside a directory, reject symbolic links, avoid mount crossings, or require a cache-only lookup. The relevant object is not the input string alone. It is the result of resolving that string against a live namespace whose directory entries, links, and mounts can change concurrently.

Software Engineering 17 Sep 2026 6 min read

Linux O_TMPFILE Keeps Staging Files Out of the Namespace

Linux O_TMPFILE creates a regular file without first placing a name for that file in a directory. The caller receives a file descriptor and can write data, set metadata, or abandon the object while no pathname exposes the partially prepared file. If publication is required, a later link operation can attach a directory entry to the same inode. This separates object construction from namespace publication. It does not make every surrounding filesystem operation transactional, and it does not provide replacement semantics for an existing destination. Its useful boundary is narrower: intermediate file state can remain reachable only through open references until the process explicitly creates a name.

Software Engineering 17 Sep 2026 6 min read

Linux mmap Keeps File Lifetime Separate From Descriptor Lifetime

A successful file-backed mmap() creates a virtual-memory mapping that does not depend on keeping the source file descriptor open. Linux explicitly permits the descriptor to be closed immediately after mmap() returns without invalidating the mapping. The mapping and the descriptor are therefore separate references with separate lifetimes. That separation is easy to miss because both originate from the same open file. It becomes operationally important when code closes descriptors aggressively, replaces pathnames, truncates files, or passes mappings across fork(). A mapped address is not a delayed read() through the original descriptor; it participates in the virtual-memory system under its own mapping contract.

Software Engineering 17 Sep 2026 6 min read

Linux inotify Events Are a Lossy Change Stream

An inotify file descriptor exposes filesystem activity as an ordered queue of event records, but that queue is not an authoritative history of namespace state. Identical unread events may be coalesced, queue capacity is bounded, and an overflow explicitly means events have been lost. A process that treats the stream as a complete transaction log can therefore preserve a state that no longer matches the filesystem. The interface is better modeled as change notification with recovery obligations. Events can make a cache current incrementally while the stream remains intact; some conditions invalidate that incremental history and require reconciliation against filesystem state.

Software Engineering 17 Sep 2026 8 min read

Linux Direct I/O Makes Alignment Part of the File Interface

Opening a regular file with O_DIRECT can make the address of a user-space buffer, the file offset, and the transfer length observable parts of the file interface. A read() or write() that is otherwise valid may fail with EINVAL when one of those values violates the direct-I/O constraints for that file. On some combinations of filesystem and kernel behavior, a misaligned operation can instead use buffered I/O. That boundary is easy to miss because ordinary buffered file I/O largely hides physical transfer geometry. The page cache and filesystem can accept an application buffer at an arbitrary address and mediate the transfer internally. Direct I/O reduces that mediation, so constraints that normally remain below the system-call boundary can become requirements on application memory and request shape.

Software Engineering 17 Sep 2026 8 min read

Linux copy_file_range Separates Copy Semantics From Copy Implementation

Linux copy_file_range Separates Copy Semantics From Copy Implementation A successful copy_file_range() call reports a byte count, not a promise about the physical path those bytes took. Linux can satisfy the request through filesystem-specific acceleration, an in-kernel transfer path, or another implementation permitted by the active filesystem interfaces. The application receives a range-copy operation with defined offset and return-value semantics; it does not receive a guarantee that storage blocks were physically duplicated.

Software Engineering 17 Sep 2026 4 min read

inotify Rename Cookies Correlate Move Events Without Making Them Atomic

A Linux rename() observed through inotify can produce two records carrying the same nonzero cookie: IN_MOVED_FROM for the old directory entry and IN_MOVED_TO for the new one. The cookie correlates those records, but it does not turn them into one atomic queue item. That boundary matters for software maintaining a pathname index, synchronizing directory state, or converting filesystem notifications into higher-level change records. A rename is one filesystem operation while its inotify representation can be a pair whose delivery has weaker grouping properties.