Skip to content

Archive

Virtual Memory

19 articles
Linux 23 Sep 2026 4 min read

MADV_DONTFORK Excludes Memory Ranges from Child Address Spaces

MADV_DONTFORK changes a specific part of Linux process creation: a mapping marked with this advice is not made available in the child created by fork(). The parent keeps the mapping. The child starts without that address range, so an address that was valid there in the parent is not automatically valid in the child. This is a semantic control over mapping inheritance, not a cache hint. It belongs to the Linux-specific madvise() operations whose effects can change memory behavior.

Tech 23 Sep 2026 5 min read

Linux fork Uses Copy-on-Write to Delay Private Page Copies

fork() creates a new process with a virtual address space derived from the caller, but Linux does not need to duplicate every private physical page at the instant the syscall returns. For ordinary private writable mappings, the kernel can arrange parent and child page tables so both processes initially refer to the same physical memory while writes are constrained by copy-on-write state. This design makes process creation proportional to page-table and kernel bookkeeping rather than to the full amount of resident private data. Physical copying is deferred until a write requires the two address spaces to diverge.

Linux 19 Sep 2026 5 min read

userfaultfd Write Protection Turns Memory Writes into Userspace Events

A thread can reach a valid, resident page and still stop before modifying it. With Linux userfaultfd write-protect mode, a registered page can be marked so that a write generates a userfaultfd page-fault event. A userspace handler receives that event, performs its bookkeeping, removes the protection, and lets the blocked thread continue. The mechanism sits between ordinary page permissions and application-level memory accounting. The page remains part of the process address space; the kernel redirects the write fault into a file-descriptor protocol instead of forcing the application to build the same control path around mprotect() and SIGSEGV.

Linux 19 Sep 2026 5 min read

userfaultfd Moves Selected Page-Fault Resolution into User Space

A memory access normally enters the kernel page-fault path and completes without an application choosing the page contents at that instant. Linux userfaultfd changes that boundary for registered virtual address ranges: selected faults become events on a file descriptor, and a user-space manager can supply or activate the page before the faulting thread continues. The mechanism does not replace the process page tables with a user-space data structure. The kernel still owns page-table state and performs the final mapping operation. User space gains control over specific fault classes and the timing of their resolution.

Tech 19 Sep 2026 6 min read

TLB Shootdowns Extend Page-Table Changes Across CPUs

TLB Shootdowns Extend Page-Table Changes Across CPUs Changing a page-table entry in memory does not by itself retire every translation derived from that entry. A CPU that previously used the mapping can retain it in a translation lookaside buffer, or TLB. On a multiprocessor system, other CPUs may hold their own cached copies, so a mapping change can require coordination beyond the CPU that modified the page table. Linux exposes this distinction through its TLB-flush interfaces. After page-table state changes, architecture code must make the affected translations unusable on every relevant CPU before software relies on the new mapping or releases memory that the old mapping could reach.

Tech 19 Sep 2026 7 min read

TLB Shootdowns Coordinate Page-Table Changes Across CPUs

A page-table entry can change in memory while another CPU still holds the old address translation in its translation lookaside buffer (TLB). Updating the page table alone therefore does not necessarily make the new mapping effective on every processor that has executed the affected address space. Operating systems close that gap with TLB invalidation. When a mapping change can make a cached translation unsafe, processors that may retain the translation must invalidate it before the kernel treats the change as globally complete. On a multiprocessor system, coordinating those remote invalidations is commonly called a TLB shootdown.

Linux 19 Sep 2026 6 min read

process_madvise Applies Memory Advice Across Process Boundaries

A process can consume memory on behalf of work that is coordinated elsewhere. Linux process_madvise() lets that external coordinator apply selected virtual-memory advice to ranges in the target process without injecting code into it. The target is identified by a pidfd, while the ranges are supplied as an array of struct iovec. That arrangement separates memory-policy decisions from the code that owns the mapping. A runtime manager, service supervisor, or memory controller can request reclaim-oriented or prefetch-oriented treatment for another process, subject to kernel support and permission checks. The system call does not transfer ownership of the mapping, freeze the target, or make its address space stable.

Linux 19 Sep 2026 6 min read

PAGEMAP_SCAN Batches Page-Table State into Address Ranges

A large virtual address range may contain only a small number of page-state transitions. Reading one pagemap entry for every virtual page exposes that state at page granularity, but it also makes user space inspect a long sequence of entries. Linux PAGEMAP_SCAN moves the filtering into the kernel and reports matching spans as struct page_region records. The interface is an ioctl() on /proc/PID/pagemap. A request supplies an address interval, category predicates, a return mask, and an output vector. The kernel walks page tables and emits contiguous regions whose selected page properties match the request. This changes the shape of page-table inspection from a stream of per-page values into a filtered range query.

Linux 19 Sep 2026 4 min read

MADV_FREE Marks Anonymous Pages for Lazy Reclaim

MADV_FREE does not immediately replace a private anonymous page with zeros. It marks eligible pages as disposable, allowing Linux to reclaim them later. Until reclaim actually occurs, existing bytes can remain observable. A write before reclaim cancels the disposable state for the affected page. That timing makes MADV_FREE distinct from advice that immediately changes the process-visible state of a range. It is a lazy reclamation contract: the application declares that old contents are expendable, while the kernel chooses when physical memory is recovered.

Linux 18 Sep 2026 5 min read

userfaultfd Turns Page Faults into Userspace Events

A thread can fault on a virtual address and remain blocked while another userspace thread or process decides what page state should make that access continue. userfaultfd provides this boundary by turning selected page faults into messages on a file descriptor and pairing those messages with ioctls that resolve the fault. The mechanism does not replace the kernel page-fault machinery. It inserts userspace control at registered ranges and fault classes, while the kernel still owns page tables, fault blocking, and the transition that makes the page usable again.

Software Engineering 18 Sep 2026 8 min read

MAP_SHARED mmap Couples Memory Writes to File-Backed Page State

A writable MAP_SHARED mapping lets a process modify file-backed state with ordinary memory stores. The bytes are addressed through virtual memory rather than passed to write(), but the mapping still participates in filesystem state: modifications can become visible through other shared mappings and file I/O, and dirty pages can later be written back to storage. That interface compresses several mechanisms into one address range. CPU stores, page faults, page-cache residency, filesystem writeback, and storage persistence can all participate in the lifetime of the same bytes. Treating a successful store as equivalent to durable file output collapses boundaries that the operating system keeps distinct.

Software Engineering 18 Sep 2026 6 min read

Linux userfaultfd Moves Selected Page Fault Handling into User Space

A page fault normally crosses from a process into the kernel and returns only after the kernel has resolved the virtual-memory condition or delivered an error. Linux userfaultfd can insert a user-space component into that path for explicitly registered address ranges. The kernel reports selected faults through a file descriptor, blocks the faulting execution context when the mode requires it, and accepts an ioctl that resolves the fault. This is a Linux virtual-memory interface, not a C or POSIX memory guarantee. Its behavior depends on negotiated kernel features, the registered range, its mapping type, and the registration mode.

Software Engineering 17 Sep 2026 6 min read

userfaultfd Moves Missing-Page Resolution into User Space

A thread can touch a valid virtual address and stop before the access completes because the page has no present backing yet. With a range registered in UFFDIO_REGISTER_MODE_MISSING, Linux can report that fault through userfaultfd instead of resolving it entirely inside the kernel. A user-space manager then decides which page contents become visible before the blocked access resumes. This changes the ownership of one part of page-fault handling. The kernel still detects the fault, validates the virtual memory area, blocks the faulting execution, and installs mappings through the UFFDIO_* interface. User space gains control over the content and timing of resolution for registered faults.

Linux 17 Sep 2026 4 min read

Truncating a Mapped File Can Trigger SIGBUS

A process can retain a valid mmap() address range after another operation shrinks the backing file, then receive SIGBUS when it touches a mapped page past the file’s new end. The mapping itself has not vanished. Its backing object no longer covers every page that the virtual mapping originally referenced. This boundary is easy to miss because mapping lifetime and file size are separate state. Closing the original file descriptor does not invalidate an established mapping, and shrinking the file does not act like munmap() on every process that maps it.

Software Engineering 17 Sep 2026 6 min read

Linux userfaultfd Turns Page Faults Into a Userspace Protocol

A thread can access a valid virtual address and stop before that access completes because another userspace component has been given responsibility for resolving the page fault. With Linux userfaultfd, selected memory ranges can turn faults into descriptor messages while the faulting thread remains blocked until an appropriate resolution operation makes progress possible. This is not a replacement for the kernel’s virtual-memory subsystem. The kernel still detects the fault, validates the registered range, blocks the affected execution path, and performs the page-table operation requested by the manager. The unusual boundary is that userspace can participate in deciding when and with what contents a fault is resolved.

Software Engineering 17 Sep 2026 6 min read

Linux mmap Keeps File Lifetime Separate From Descriptor Lifetime

A successful file-backed mmap() creates a virtual-memory mapping that does not depend on keeping the source file descriptor open. Linux explicitly permits the descriptor to be closed immediately after mmap() returns without invalidating the mapping. The mapping and the descriptor are therefore separate references with separate lifetimes. That separation is easy to miss because both originate from the same open file. It becomes operationally important when code closes descriptors aggressively, replaces pathnames, truncates files, or passes mappings across fork(). A mapped address is not a delayed read() through the original descriptor; it participates in the virtual-memory system under its own mapping contract.

Software Engineering 17 Sep 2026 8 min read

File-Backed Mappings Can Outlive the File Size They Assume

A process can retain a valid virtual memory mapping after another actor has shortened the mapped file. The address range still exists in the process, but the backing object may no longer contain every page that range once represented. On POSIX systems, an access to a whole mapped page beyond the new end can deliver SIGBUS rather than behaving like an ordinary failed file read. That boundary makes file-backed mmap() different from copying bytes into private heap storage. A pointer into a mapping is not proof that the corresponding file extent still exists. The virtual address, mapping lifetime, file identity, and current file size are related state, but they are not one indivisible object.

Tech 16 Sep 2026 7 min read

TLB Caches Recent Virtual Address Translations

Modern processors commonly execute programs in virtual address spaces. A load or store can begin with a virtual address while the memory system ultimately needs a physical location and access permissions. Page tables hold the mapping information, but consulting their hierarchy for every memory reference would add substantial work. A translation lookaside buffer, or TLB, keeps recently used address translations near the processor. A TLB hit supplies cached mapping information without a full page-table walk. A TLB miss triggers additional translation work even when the requested application data is already present in a CPU cache.

Tech 03 Sep 2026 9 min read

What Virtual Memory Does When Your Computer Runs Low on RAM

You can sometimes open more applications than seem able to fit in your computer’s physical memory. At other times, opening one more browser tab makes the whole machine feel sluggish even though nothing has crashed. Virtual memory helps explain both situations. It lets the operating system manage memory without requiring every piece of an application’s active data to remain in physical RAM at the same time. When RAM becomes scarce, the system can reclaim space in several ways, including moving some memory contents to storage.