Skip to content

Archive

Memory

37 articles
Artificial Intelligence 24 Sep 2026 4 min read

Activation Checkpointing Trades Saved Activations for Recomputation

Activation checkpointing changes which forward-pass tensors remain resident until backpropagation. Instead of retaining every intermediate activation required by gradient computation, a checkpointed region keeps selected boundary state and reconstructs discarded intermediates when the backward pass reaches that region. The mechanism reduces activation memory at the cost of extra computation. It does not shrink model parameters, optimizer state, or gradients, so its effect on total training memory depends on how much of the footprint comes from activations.

Linux 23 Sep 2026 3 min read

MADV_WIPEONFORK Replaces Inherited Private Memory with Zeroes

MADV_WIPEONFORK Replaces Inherited Private Memory with Zeroes A normal fork() gives the child mappings derived from the parent’s address space, with private writable pages commonly handled through copy-on-write. MADV_WIPEONFORK changes that inheritance rule for a selected private anonymous range: the mapping remains present in the child, but its contents are zero-filled there. The parent keeps its existing bytes. The operation therefore changes child-visible memory at the process-creation boundary rather than erasing the parent’s range.

Tech 22 Sep 2026 7 min read

NUMA First-Touch Placement Ties Physical Pages to the Faulting CPU

On a NUMA machine, reserving virtual address space does not necessarily choose the physical NUMA node that will supply every page. For anonymous memory, physical allocation commonly happens later, when a CPU first faults on a page. Under the default local allocation policy, that fault can make initialization order part of the memory-placement decision. This behavior is often called first-touch placement. The thread that first writes or otherwise faults a page can cause Linux to allocate backing memory near the NUMA node on which that thread is running, subject to the active memory policy, cpuset constraints, available memory, and kernel fallback behavior.

Tech 22 Sep 2026 4 min read

io_uring Registered Buffers Pin User Pages for Repeated I/O

Linux io_uring can submit asynchronous I/O with ordinary user buffers, but repeated operations may still require the kernel to resolve and pin the relevant user pages for each request. Registered buffers move part of that work into an explicit setup phase. An application registers one or more memory regions with the ring. The kernel records those regions and keeps the backing pages pinned while the registration remains active. Later requests can refer to a registered region by index instead of presenting an arbitrary buffer that must be prepared from scratch.

Linux 21 Sep 2026 6 min read

memfd File Seals Freeze Shared Memory State

A memfd_create() file starts as a mutable anonymous file. It can be resized, written, and mapped much like a regular file, while its storage remains volatile and disappears after the last reference is released. File sealing adds a different phase to that lifecycle: after data has been populated, the kernel can permanently reject selected classes of later modification. That transition is useful when one process prepares bytes and then hands the same file description to another process. The receiver can inspect the seals attached to the inode instead of relying only on a convention that the sender will stop changing the object.

Software Engineering 19 Sep 2026 5 min read

process_vm_readv and process_vm_writev Transfer Memory Across Process Boundaries

process_vm_readv() and process_vm_writev() let one Linux process copy bytes directly between its address space and another process’s address space. The calls operate on vectors of local and remote memory ranges, but a successful process lookup does not make remote memory stable. Mapping changes, page accessibility, permissions, and concurrent mutation remain separate parts of the contract. These interfaces are Linux-specific system calls. They do not define C object lifetime, synchronization, or a portable interprocess-memory model.

Artificial Intelligence 19 Sep 2026 6 min read

Ollama Gemma 3 270M Memory Is More Than the Model File

Ollama lists gemma3:270m at about 292 MB. That number is useful for storage planning, but it is not a RAM requirement. It describes the packaged model data for the default Ollama variant, which uses Q8_0 quantization. Once inference starts, the runtime also needs memory for model metadata, execution buffers, token state, and the key-value cache used by attention. That distinction matters on small machines. A device with 512 MB of RAM may appear large enough when compared only with a 292 MB model file, yet the remaining memory must also accommodate Ollama and the operating system. The context configuration can move the total substantially.

Linux 19 Sep 2026 4 min read

memfd Seals Constrain Shared-Memory Mutation After Handoff

A memfd_create() descriptor names an anonymous file whose storage lives in memory-backed filesystem infrastructure. By itself, descriptor handoff does not freeze that object: a process retaining suitable access can still write bytes, truncate the file, or extend it. Linux file seals add kernel-enforced restrictions that can make selected mutations fail after the producer declares the object complete. This changes shared-memory handoff from a convention into a state transition enforced at the file object.

Linux 19 Sep 2026 4 min read

MADV_FREE Marks Anonymous Pages for Lazy Reclaim

MADV_FREE does not immediately replace a private anonymous page with zeros. It marks eligible pages as disposable, allowing Linux to reclaim them later. Until reclaim actually occurs, existing bytes can remain observable. A write before reclaim cancels the disposable state for the affected page. That timing makes MADV_FREE distinct from advice that immediately changes the process-visible state of a range. It is a lazy reclamation contract: the application declares that old contents are expendable, while the kernel chooses when physical memory is recovered.

Tech 19 Sep 2026 6 min read

IOMMU Translation Separates Device DMA Addresses from Physical Memory

A DMA-capable device can issue memory transactions without the CPU copying each payload. On systems with an IOMMU, the address carried by that device transaction does not have to be a host physical address. The IOMMU can translate a device-visible I/O virtual address into a physical page and reject accesses outside the configured mapping. That translation boundary changes both isolation and data movement. Drivers and operating systems can give a device a constrained address space, while the hardware must maintain translation state close enough to the I/O path to avoid turning every DMA request into a page-table walk.

Tech 19 Sep 2026 5 min read

Cache-Line False Sharing Moves Coherence Ownership Between CPUs

Cache-Line False Sharing Moves Coherence Ownership Between CPUs Two threads can update different variables without sharing a lock or touching the same bytes and still interfere at the hardware level. If those variables occupy the same cache line, a coherent multiprocessor treats their storage as one coherence unit. Repeated writes from different CPUs can therefore move ownership of that line between caches even though the program considers the variables independent.

Software Engineering 18 Sep 2026 4 min read

memfd Seals Turn Shared File State into Monotonic Restrictions

A memfd_create() file can begin as writable shared state and later become progressively more constrained. File seals make that transition monotonic: successful seals are properties of the inode, affect every descriptor referring to it, and cannot be removed. That property is useful when one process prepares bytes and then transfers a descriptor to another process. The receiver can inspect kernel-enforced restrictions instead of relying only on a protocol promise that the producer has stopped changing the object.

Tech 17 Sep 2026 6 min read

Write Combining Merges Adjacent Stores Before Memory Traffic

Some memory regions are written far more often than they are read. Frame buffers, device apertures, and streaming output areas are common examples. Sending every small CPU store as a separate memory transaction can waste bus bandwidth and transaction overhead. Write combining gives the processor a temporary place to collect compatible stores. Several writes targeting nearby addresses can be merged into a larger transaction before they leave the CPU. The technique favors sustained write throughput, but it changes the timing and ordering properties that software can safely assume.

Tech 17 Sep 2026 6 min read

NUMA Makes Memory Location Part of Access Cost

NUMA Makes Memory Location Part of Access Cost A large multiprocessor server can expose one physical address space while giving different processors different paths to that memory. A load from a page attached to the processor running a thread can take a shorter route than a load from memory attached to another processor package or NUMA node. This arrangement is called non-uniform memory access, or NUMA. It lets systems scale memory capacity and bandwidth across multiple processor sockets or chiplet groups without forcing every memory request through one centralized controller.

Software Engineering 17 Sep 2026 4 min read

memfd Seals Turn Shared Files into Monotonic Objects

A Linux memfd can begin as a writable anonymous file and later acquire restrictions that cannot be removed. The restrictions belong to the inode, so transferring or duplicating a descriptor does not create a less restricted view. Once a seal is added successfully, every descriptor referring to that inode is subject to it. This makes sealing different from descriptor access modes. A descriptor can carry local flags, while a seal changes the mutation boundary of the shared file object itself.

Tech 17 Sep 2026 7 min read

IOMMU Remaps Device DMA Addresses for Memory Isolation

Direct memory access lets a device move data between itself and system memory without making the CPU copy every byte. Network adapters, storage controllers, GPUs, and other high-throughput devices rely on DMA to keep data moving efficiently. That capability also creates a protection problem. A device that can issue unrestricted memory transactions could read or overwrite physical pages belonging to the kernel, another process, or another virtual machine. An input-output memory management unit, commonly called an IOMMU, places address translation and access control between DMA-capable devices and physical memory.

Tech 16 Sep 2026 7 min read

TLB Caches Recent Virtual Address Translations

Modern processors commonly execute programs in virtual address spaces. A load or store can begin with a virtual address while the memory system ultimately needs a physical location and access permissions. Page tables hold the mapping information, but consulting their hierarchy for every memory reference would add substantial work. A translation lookaside buffer, or TLB, keeps recently used address translations near the processor. A TLB hit supplies cached mapping information without a full page-table walk. A TLB miss triggers additional translation work even when the requested application data is already present in a CPU cache.

Linux 16 Sep 2026 6 min read

Linux Readahead Expands Sequential Page-Cache Reads

A buffered file read can cause Linux to fetch more data than the application explicitly requested. The extra I/O is readahead: the kernel populates nearby page-cache folios in anticipation of continued access. This behavior sits between application read size and storage request size. A process may issue modest read() calls while the kernel submits larger reads to keep later accesses from waiting on storage. Readahead is page-cache speculation Buffered file I/O normally passes through the page cache. When requested file data is absent, the kernel must arrange I/O for that miss. The readahead path can extend that operation across additional folios that are not yet present in the cache.

Linux 16 Sep 2026 6 min read

Linux cgroup memory.high Converts Overage into Reclaim Pressure

A cgroup can remain alive after its memory usage crosses memory.high. The boundary does not behave like a hard allocation ceiling: tasks in the cgroup are throttled and pushed into heavy reclaim pressure, and usage can remain above the configured value under extreme conditions. That behavior makes memory.high materially different from memory.max. The former converts excess usage into execution cost and reclaim work. The latter is a hard limit that can lead to a cgroup OOM when reclaim cannot reduce usage enough.

Go 16 Sep 2026 5 min read

Go sync.Pool Items Can Disappear Across Garbage Collection

A value placed in a Go sync.Pool is not guaranteed to remain there until a later Get. The runtime may remove pooled items automatically, so the pool acts as a reuse opportunity rather than durable storage. That property shapes both the performance profile and the correctness boundary of sync.Pool. Code can benefit when an object survives long enough to be reused, but it must remain correct when every Get behaves as if no prior item were available.

Tech 16 Sep 2026 6 min read

DRAM Refresh Restores Charge Before Bits Fade

Dynamic random-access memory stores data in cells whose electrical state does not remain stable indefinitely. Charge leaks from a cell over time, even when software performs no reads or writes. A memory system therefore has to refresh DRAM periodically to preserve stored bits. Refresh is a maintenance operation rather than a request from an application. The memory controller and DRAM device coordinate it alongside ordinary reads and writes. During parts of that work, some memory resources cannot serve normal requests.

Tech 16 Sep 2026 6 min read

CPU Store Buffers Decouple Retirement from Cache Writes

A processor does not need every store instruction to finish its cache update before later instructions make progress. Modern cores commonly place completed stores into a store buffer, allowing the instruction to retire while the memory subsystem handles the write afterward. This separation improves throughput because cache ownership, coherence traffic, and other memory activity can take longer than the execution pipeline can afford to wait. The buffer acts as a queue between architectural execution and the cache hierarchy.

Tech 16 Sep 2026 6 min read

CPU Cache Associativity Limits Where Lines Can Reside

Processor caches keep recently used memory close to execution cores, but cache capacity alone does not determine which data can remain resident. Most general-purpose CPU caches divide storage into sets and give each set a fixed number of slots, commonly called ways. A memory block maps to a particular set. It can occupy any way inside that set, but it cannot move into an unrelated set merely because that other set has free space. This placement rule makes hardware lookup practical and fast, while creating a distinct source of misses when too many active blocks compete for the same set.

Artificial Intelligence 14 Sep 2026 7 min read

Trade Activation Memory for Recomputation with Checkpointing

Backpropagation needs intermediate values from the forward pass to compute parameter and input gradients. Keeping every required activation alive until its gradient is calculated can consume substantial device memory, especially as model depth, batch size, or sequence length grows. Activation checkpointing changes which intermediates are retained. Selected boundary tensors remain available, while activations inside a checkpointed region are discarded after the forward pass and produced again when the backward pass reaches that region. The model computes the same conceptual function, but the execution schedule exchanges additional computation for lower activation storage.