Skip to content

Archive

Linux

231 articles
Tech 19 Sep 2026 6 min read

PCIe Posted MMIO Writes Can Outlive the CPU Store That Issued Them

A CPU can retire or complete an MMIO store before the corresponding PCIe Memory Write has reached the target device. The gap exists because PCIe Memory Write requests are posted: the requester sends them without waiting for a completion packet from the completer. That property is useful for throughput, but it creates an important boundary. A software-visible store instruction, an ordering barrier, and device observation of the write are not automatically the same event.

Linux 19 Sep 2026 6 min read

PAGEMAP_SCAN Batches Page-Table State into Address Ranges

A large virtual address range may contain only a small number of page-state transitions. Reading one pagemap entry for every virtual page exposes that state at page granularity, but it also makes user space inspect a long sequence of entries. Linux PAGEMAP_SCAN moves the filtering into the kernel and reports matching spans as struct page_region records. The interface is an ioctl() on /proc/PID/pagemap. A request supplies an address interval, category predicates, a return mask, and an output vector. The kernel walks page tables and emits contiguous regions whose selected page properties match the request. This changes the shape of page-table inspection from a stream of per-page values into a filtered range query.

Linux 19 Sep 2026 6 min read

openat2 Constrains Path Resolution at the Kernel Boundary

A pathname that begins inside a trusted directory can resolve somewhere else before open() returns. Parent components, symbolic links, magic links, mount points, and concurrent namespace changes all participate in Linux pathname lookup. Checking a string before opening it therefore does not establish where the kernel will finish resolution. Linux openat2() places restrictions inside the lookup operation itself. A caller supplies a directory file descriptor, ordinary open flags, and a resolve policy in struct open_how. The kernel then applies those constraints while walking every relevant path component. This moves a security boundary from pre-validation of pathname text into the operation that actually resolves the pathname.

Linux 19 Sep 2026 5 min read

openat2 Constrains Linux Path Resolution at the Open Boundary

A pathname can change meaning while a process is resolving it. Directory renames, symbolic links, mount points, and .. components can redirect lookup away from the directory a program intended to treat as its boundary. Linux openat2() attaches resolution policy to the lookup itself. Its struct open_how contains a resolve bit mask, so the kernel can reject a path when resolution violates a caller-selected constraint instead of relying only on checks performed before open().

Software Engineering 19 Sep 2026 8 min read

Open File Description Locks Bind Byte Ranges to File Instances

A byte-range lock can protect the same inode yet have radically different lifetime semantics depending on what owns the lock. Traditional fcntl() record locks are process-associated. Open file description locks instead attach to the kernel open file description referenced by a descriptor. That shift changes which close operation releases a lock, what survives fork(), and whether two threads in one process can contend on the same file region. Linux exposes this model through F_OFD_SETLK, F_OFD_SETLKW, and F_OFD_GETLK. The range model remains familiar: struct flock specifies a read lock, write lock, or unlock together with an offset and length. The significant difference is ownership.

Cybersecurity 19 Sep 2026 5 min read

no_new_privs Blocks Exec-Time Privilege Gain

A service may need to execute helper programs after it has accepted untrusted input. If one of those programs is set-user-ID, set-group-ID, or carries file capabilities, a normal execve() can cross a privilege boundary even when the calling process itself has no intent to acquire extra authority. Linux no_new_privs changes that transition: once set for a thread, later execve() calls cannot grant privileges that were absent from the caller at the point of execution.

Tech 19 Sep 2026 5 min read

NIC Interrupt Moderation Trades Wakeup Rate for Packet Latency

A network adapter does not need to interrupt a CPU for every received packet or completed transmission. Many NICs can hold interrupt delivery briefly and report several completion events together. This interrupt moderation reduces interrupt traffic and CPU entry overhead, but it can also delay the moment software notices newly completed work. The mechanism sits between packet DMA and the driver’s receive or transmit processing. It changes notification timing; it does not change the packet’s wire format, Ethernet ordering rules, or the basic requirement that the driver eventually process completed descriptors.

Cybersecurity 19 Sep 2026 6 min read

memfd Seals Turn Shared Memory into a Kernel-Enforced Immutable Payload

Shared memory is efficient partly because two processes can observe the same storage without copying it. That property becomes a security problem when one side validates bytes and later consumes them while another side still holds authority to mutate the same object. Linux memfd sealing can narrow that race by making selected mutations fail in the kernel before the file descriptor crosses a trust boundary. memfd_create() creates an anonymous file and returns an ordinary file descriptor. The object can be sized, written, mapped, and transferred over a UNIX domain socket. With MFD_ALLOW_SEALING, the inode starts with an empty seal set, allowing the producer to add irreversible restrictions after population.

Linux 19 Sep 2026 4 min read

memfd Seals Constrain Shared-Memory Mutation After Handoff

A memfd_create() descriptor names an anonymous file whose storage lives in memory-backed filesystem infrastructure. By itself, descriptor handoff does not freeze that object: a process retaining suitable access can still write bytes, truncate the file, or extend it. Linux file seals add kernel-enforced restrictions that can make selected mutations fail after the producer declares the object complete. This changes shared-memory handoff from a convention into a state transition enforced at the file object.

Linux 19 Sep 2026 4 min read

MADV_FREE Marks Anonymous Pages for Lazy Reclaim

MADV_FREE does not immediately replace a private anonymous page with zeros. It marks eligible pages as disposable, allowing Linux to reclaim them later. Until reclaim actually occurs, existing bytes can remain observable. A write before reclaim cancels the disposable state for the affected page. That timing makes MADV_FREE distinct from advice that immediately changes the process-visible state of a range. It is a lazy reclamation contract: the application declares that old contents are expendable, while the kernel chooses when physical memory is recovered.

Software Engineering 19 Sep 2026 6 min read

Linux pidfds Bind Process Operations to Stable Kernel References

A numeric process ID names a process only while that PID remains assigned to it. After process exit and reaping, Linux may reuse the number for another process. Code that observes a PID, performs unrelated work, then acts on that number can therefore cross a lifetime boundary that the integer itself does not encode. Linux pidfds provide a file-descriptor reference to a process so later operations can target the referenced process object rather than repeat a numeric PID lookup.

Software Engineering 19 Sep 2026 6 min read

Linux membarrier Moves Memory Ordering Cost to a Coordinating Thread

A concurrent runtime can have thousands of fast-path operations for every rare state transition that requires global coordination. Placing a full memory barrier on every fast path makes each operation pay for that rare transition. Linux membarrier() supports the opposite arrangement: a coordinating thread enters the kernel and forces a defined ordering point across a target set of threads, moving more cost to the infrequent side of the protocol. This is not a generic replacement for atomics, mutexes, or language memory models. It is a Linux kernel interface whose guarantees apply to memory accesses and targeted threads under specific commands. Correct use requires a protocol that already defines which accesses occur before and after the coordination point.

Linux 19 Sep 2026 5 min read

io_uring Multishot Requests Persist Across Completion Events

A normal io_uring request has a simple lifetime: userspace submits one SQE and eventually receives one CQE. Multishot operations change that relationship. One submitted request can remain active in the kernel and produce several completion queue entries as matching events occur. That persistence changes completion handling from a one-CQE-per-request assumption into an explicit lifecycle protocol. The decisive state is carried by IORING_CQE_F_MORE: when the flag is present, the originating request can produce another completion; when it is absent, that multishot request has terminated.

Software Engineering 19 Sep 2026 6 min read

io_uring Linked Requests Encode Dependency in Submission Order

An io_uring submission queue can contain many operations at once, but not every operation has to be independent. Setting IOSQE_IO_LINK on a submission queue entry binds it to the next entry, forming a chain in which execution order and failure propagation become part of the kernel-visible request structure. That changes the contract compared with submitting two unrelated SQEs and coordinating them after completion. A linked chain expresses dependency before the kernel starts processing the operations. The distinction matters when a later request is valid only after an earlier request has completed, or when failure of one stage should prevent the remaining stages from running.

Cybersecurity 19 Sep 2026 6 min read

IMA Appraisal Binds File Access to Integrity Metadata

A file can have ordinary read or execute permission and still fail an integrity check at the kernel boundary. Linux Integrity Measurement Architecture (IMA) appraisal can apply policy rules at selected hooks and require file content to agree with integrity metadata before the covered operation proceeds. This adds a content-integrity condition to access decisions without turning Unix mode bits or application authorization into integrity mechanisms. IMA contains related but distinct functions. Measurement records file state in an integrity measurement list and can extend measurements into a TPM. Appraisal evaluates a file against integrity metadata and can reject access when policy and enforcement mode require a valid result. Audit records security-relevant state. A deployment that only measures files gains evidence about observed state; it does not automatically gain the blocking behavior associated with appraisal.

Linux 19 Sep 2026 6 min read

Idmapped Mounts Remap Ownership Without Rewriting Inodes

The same inode can appear with different ownership through two mount points without any recursive chown(). Linux idmapped mounts attach an ID mapping to a mount, so VFS ownership presentation and permission checks can translate user and group IDs for that view while the ownership stored by the filesystem remains unchanged. This property separates persistent inode metadata from the identity view exposed at a particular mount. It is especially useful when a filesystem tree must be shared with a container whose user namespace maps IDs differently from the host.

Linux 19 Sep 2026 4 min read

How Linux Maps an ESP32 USB-UART Bridge to /dev/ttyUSB0 on Fedora

An ESP32 development board connected through a CH341 USB-to-UART bridge does not appear on Fedora as a Windows-style COM port. Linux binds the USB interface to a serial driver and exposes a character device such as /dev/ttyUSB0. A working connection can produce this kernel message: usb 5-1: ch341-uart converter now attached to ttyUSB0 That line confirms USB enumeration, binding to the ch341 serial driver, and creation of ttyUSB0. The path used by a flasher or serial monitor is therefore /dev/ttyUSB0.

Cybersecurity 19 Sep 2026 6 min read

fscrypt Policies Bind Directory Trees to Filesystem Encryption Keys

A directory can remain fully visible in a mounted Linux filesystem while its regular-file contents and filenames are unusable without a particular key. With fscrypt, that boundary is attached to filesystem objects rather than created by mounting a second encrypted filesystem. An encryption policy assigned to an empty directory is inherited by regular files, directories, and symbolic links created beneath it. The property is narrower than full filesystem secrecy. fscrypt encrypts file contents and filenames, but most filesystem metadata remains visible, and ordinary permission checks continue to define who may access objects once the relevant key is present. Encryption policy and access control are therefore separate boundaries.

Cybersecurity 19 Sep 2026 7 min read

fs-verity Binds Read-Only File Reads to a Merkle-Tree Digest

A file can be stored on media that is less trusted than the process consuming it. Making that file read-only through ordinary permission bits does not prove that the bytes later returned from storage are the bytes that were approved earlier. Linux fs-verity addresses that narrower integrity boundary for supported filesystems by binding reads from an enabled file to a Merkle tree and a stable file digest. The mechanism has two distinct security roles. The kernel verifies file data against the Merkle tree during reads. A separate policy must establish that the resulting fs-verity digest is the digest that the system intended to trust. Treating those roles as one guarantee overstates what the filesystem feature provides.

Cybersecurity 19 Sep 2026 7 min read

fanotify Permission Events Put File Access Behind a User-Space Decision

A process can pass ordinary filesystem permission checks and still wait before its file operation completes. Linux fanotify permission events let a monitoring group intercept selected operations and require a user-space listener to return an allow or deny response. The mechanism inserts a synchronous decision point into the access path rather than merely reporting activity after it occurs. That distinction makes fanotify useful for security products that need content inspection or policy evaluation close to file access. It also creates a dependency that ordinary notification systems do not have: the kernel may be holding another process at a permission event while user space decides its fate.

Linux 19 Sep 2026 5 min read

eventfd Turns Kernel Notifications into Pollable Counters

A Linux process can signal work through a file descriptor without moving a byte stream between producer and consumer. eventfd() creates a kernel-maintained 64-bit counter whose readiness can be observed by poll(), select(), or epoll. A write adds to the counter; a read consumes its accumulated state according to the descriptor mode. That shape makes eventfd different from a pipe. A pipe preserves a sequence of bytes. An eventfd preserves counter state. When the application needs a wakeup edge plus a compact amount of accumulated state, that distinction removes buffering and framing that a byte stream would otherwise require.

Software Engineering 19 Sep 2026 5 min read

EPOLLET Makes Readiness a State-Transition Contract

A descriptor registered with EPOLLET can remain readable after an event has been delivered without appearing again in the next epoll_wait(). The kernel reports a readiness transition; it does not promise to repeat the same notification merely because unread data remains. That distinction turns edge-triggered epoll into a state-transition contract between the kernel and the event loop. The consequence is structural. A handler cannot treat one event as permission for one read() and then return to the wait loop. With edge-triggered monitoring, the handler must account for all immediately available I/O state before relying on another transition.

Software Engineering 19 Sep 2026 5 min read

Edge-Triggered epoll Requires Draining Readiness to EAGAIN

With EPOLLET, an event loop can consume one notification, read only part of the available data, and then wait indefinitely even though unread bytes remain in the socket buffer. The descriptor is still ready, but no new readiness transition has occurred to generate another edge. That behavior makes edge-triggered epoll a contract between notification semantics and nonblocking I/O. The event says that readiness changed; it is not a promise that the kernel will keep repeating the same notification until the application finishes the work.

Linux 19 Sep 2026 5 min read

Configure Wi-Fi on Ubuntu Server Without Guessing the Netplan Backend

A Netplan Wi-Fi block can be syntactically valid and still be wrong for a particular Ubuntu Server. The two values that cannot safely be copied from an example are the interface name and the renderer. A configuration that names wlp2s0 assumes the machine actually has an interface with that name. Setting renderer: networkd assumes the installation is intended to use systemd-networkd; for Wi-Fi, that backend also relies on wpa_supplicant. Other installations may already be managed by NetworkManager.