Skip to content

Archive

Storage

64 articles
Software Engineering 22 Sep 2026 6 min read

Tombstones Prevent Deleted Data from Reappearing

Tombstones Prevent Deleted Data from Reappearing Deletion is not merely the absence of a value in a replicated store. Absence carries no information about whether a key was deliberately removed or whether a replica has simply never received it. When replicas can be temporarily disconnected, that distinction determines whether synchronization preserves a deletion or accidentally restores old data. A tombstone records the deletion as versioned state. Replicas can compare that marker with older values and keep the deletion when they reconcile. The marker can eventually be reclaimed, but only after the system has a defensible boundary beyond which an older value cannot return.

Tech 22 Sep 2026 7 min read

NVMe Completion Queue Phase Tags Mark Reused Entries

An NVMe completion queue is a fixed-size circular array in host memory. The controller writes completion queue entries as commands finish, while host software consumes those entries and advances the queue head. Eventually both sides return to slots that already contain data from an earlier circuit of the ring. Reusing memory creates a small but important ambiguity. A slot can contain a perfectly formed completion entry even when the controller has not written a new completion there yet. Clearing every consumed entry would add memory traffic and still require careful coordination between the host and controller.

Tech 22 Sep 2026 6 min read

NVMe Completion Queue Phase Tags Distinguish New Entries After Ring Wrap

An NVMe completion queue is a circular memory structure shared by a controller and host software. The controller posts completion queue entries after commands finish, while the host consumes those entries and advances its queue head. Once either side reaches the final slot, its index wraps to slot zero. That wrap creates a small but important state problem. Queue memory still contains bytes from earlier completions. Reading a nonzero entry at the current head is not enough to prove that the controller has posted a fresh completion there. NVMe solves this with a one-bit Phase Tag carried in every completion queue entry.

Tech 22 Sep 2026 5 min read

Linux Buffered Writes Separate Write Completion from Storage Persistence

A successful buffered write() does not generally mean that the new file data has already reached non-volatile storage. On Linux, the common buffered I/O path places file data in the page cache, marks the affected cache state dirty, and lets storage I/O occur later. That separation is central to normal filesystem I/O. Memory absorbs application writes at CPU-accessible speed, while the kernel can schedule backing-device traffic independently. The result improves flexibility and can reduce immediate storage stalls, but it also creates a boundary between syscall completion and persistence.

Tech 22 Sep 2026 4 min read

io_uring Registered Buffers Pin User Pages for Repeated I/O

Linux io_uring can submit asynchronous I/O with ordinary user buffers, but repeated operations may still require the kernel to resolve and pin the relevant user pages for each request. Registered buffers move part of that work into an explicit setup phase. An application registers one or more memory regions with the ring. The kernel records those regions and keeps the backing pages pinned while the registration remains active. Later requests can refer to a registered region by index instead of presenting an arbitrary buffer that must be prepared from scratch.

Tech 17 Sep 2026 9 min read

NVMe Queue Pairs Separate Command Submission from Completion

NVMe Queue Pairs Separate Command Submission from Completion An NVMe solid-state drive does not need the CPU to hand each storage command directly to a device register and then wait for that command to finish. Instead, NVMe places command and completion records in queues held in host memory. The controller reads pending commands from submission queues and writes results to associated completion queues. That arrangement matches fast PCIe storage well. Modern SSD controllers can process many operations at once across flash channels, internal dies, and controller pipelines. A queue model lets software keep that parallel hardware busy while avoiding a long series of synchronous command handoffs.

Software Engineering 17 Sep 2026 8 min read

Linux Direct I/O Makes Alignment Part of the File Interface

Opening a regular file with O_DIRECT can make the address of a user-space buffer, the file offset, and the transfer length observable parts of the file interface. A read() or write() that is otherwise valid may fail with EINVAL when one of those values violates the direct-I/O constraints for that file. On some combinations of filesystem and kernel behavior, a misaligned operation can instead use buffered I/O. That boundary is easy to miss because ordinary buffered file I/O largely hides physical transfer geometry. The page cache and filesystem can accept an application buffer at an arbitrary address and mediate the transfer internally. Direct I/O reduces that mediation, so constraints that normally remain below the system-call boundary can become requirements on application memory and request shape.

Tech 16 Sep 2026 5 min read

SSD TRIM Marks Unused Data for Controller Reclamation

Deleting a file changes filesystem metadata, but an SSD does not automatically know that every flash page formerly associated with the file can be treated as disposable. From the drive’s point of view, previously written logical block addresses can remain valid until the host explicitly indicates otherwise or overwrites them. TRIM provides that indication for ATA storage. Comparable deallocation commands exist in other storage protocols. The host identifies logical block ranges whose previous contents no longer need to be preserved, and the SSD can use that information when managing flash internally.

Tech 16 Sep 2026 5 min read

SSD Garbage Collection Amplifies Host Writes

An SSD can write substantially more data to NAND than the host sends to the device. The extra traffic appears when the controller must relocate still-valid pages before reclaiming flash blocks that contain invalid data. This internal movement is write amplification. It is a consequence of the mismatch between fine-grained logical updates and NAND erase constraints, not an extra write issued by the application. NAND pages cannot be overwritten in place NAND flash is programmed in pages but erased in larger erase blocks. A page that already contains programmed data cannot generally receive an arbitrary in-place replacement. The controller writes the new version to another available page and marks the old physical page as stale in its mapping state.

Tech 16 Sep 2026 7 min read

SATA Native Command Queuing Reorders Storage Requests

A storage request does not always need to finish in the same order that software submitted it. SATA Native Command Queuing, commonly shortened to NCQ, lets a compatible host issue several commands without waiting for each one to complete first. The device can then schedule eligible work in an order suited to its internal operation. The mechanism was especially valuable for hard disk drives, where physical head movement and rotational position can make request order affect service time. Solid-state drives have no moving heads, but multiple outstanding commands can still expose parallel work to the controller and reduce idle gaps.

Database 16 Sep 2026 5 min read

PostgreSQL VACUUM Tail Truncation Requires an Exclusive Lock

Plain PostgreSQL VACUUM normally leaves reclaimed heap space inside the relation for later reuse. A distinct tail-truncation phase can instead shorten the relation file when a contiguous run of empty pages exists at its physical end. That phase requires an ACCESS EXCLUSIVE lock. The lock boundary makes tail truncation materially different from ordinary vacuum cleanup. Routine heap and index maintenance is designed to coexist with normal reads and writes, while shortening the physical relation requires a brief period in which concurrent table access cannot proceed.

Database 16 Sep 2026 5 min read

PostgreSQL HOT Updates Avoid New Index Entries

A PostgreSQL UPDATE can create a new physical row version without creating corresponding new entries in ordinary tuple-addressing indexes. This heap-only tuple optimization, commonly called HOT, applies when the replacement tuple remains on the same heap page and the update does not change values that disqualify HOT for the table’s indexes. That distinction matters because PostgreSQL implements MVCC updates by retaining row versions rather than overwriting a tuple in place. Without HOT, an update can add work to both the heap and every index even when the indexed key values remain stable.

Database 16 Sep 2026 5 min read

PostgreSQL B-Tree Deduplication Compresses Duplicate Keys

A PostgreSQL B-tree leaf page can hold many index tuples with identical key values. When deduplication is applicable, PostgreSQL can represent a group of those tuples as one posting-list tuple: the indexed key appears once, followed by a sorted array of heap tuple identifiers. This representation changes physical index density without changing the logical set of index entries. Each heap tuple remains individually addressable through its TID, while repeated key material occupies less leaf-page space.

Tech 16 Sep 2026 3 min read

NVMe Doorbell Registers Notify Controllers of Queue Progress

NVMe Doorbell Registers Notify Controllers of Queue Progress NVMe places submission and completion queues in host memory, but a controller still needs a signal when software adds commands or consumes completion entries. Doorbell registers provide that signal. Host software writes queue pointer values to memory-mapped controller registers so the device can track progress without scanning host memory continuously. The mechanism separates queue storage from queue notification. Commands and completion entries live in DMA-accessible memory, while small register writes tell the controller which portion of each queue has changed.

Linux 16 Sep 2026 6 min read

Linux Readahead Expands Sequential Page-Cache Reads

A buffered file read can cause Linux to fetch more data than the application explicitly requested. The extra I/O is readahead: the kernel populates nearby page-cache folios in anticipation of continued access. This behavior sits between application read size and storage request size. A process may issue modest read() calls while the kernel submits larger reads to keep later accesses from waiting on storage. Readahead is page-cache speculation Buffered file I/O normally passes through the page cache. When requested file data is absent, the kernel must arrange I/O for that miss. The readahead path can extend that operation across additional folios that are not yet present in the cache.

Tech 16 Sep 2026 5 min read

fsync on a File Does Not Persist Its Directory Entry

A successful fsync() on a regular file does not, by itself, guarantee that the directory entry naming that file has reached persistent storage. Linux documents this boundary explicitly: file synchronization covers the file’s data and associated metadata, while persistence of the containing directory entry requires an fsync() on a file descriptor for that directory. This distinction matters when software creates a new file or atomically replaces an existing pathname. File contents and pathname metadata are separate pieces of filesystem state, and a crash can test the durability boundary between them.

Tech 15 Sep 2026 5 min read

SSD Write Cache and Sustained Transfer Speed

An SSD can copy the first part of a large file at high speed, then settle at a much lower rate even though nothing else appears to have changed. That drop can be normal. Many consumer SSDs use part of their NAND as a fast write cache, allowing short bursts to finish before the drive has to sustain writes in its denser storage mode. This behavior makes a single peak transfer number a poor description of every write workload. Cache size, free space, NAND type, controller policy, temperature, and the amount of data already waiting inside the drive can all affect the speed seen during a long transfer.

Tech 15 Sep 2026 6 min read

SSD TRIM Marks Discarded Data for Flash Reuse

Deleting a file changes filesystem metadata, but that action does not automatically tell a solid-state drive which flash pages no longer contain useful data. From the drive’s point of view, previously written logical block addresses can remain valid until the host explicitly replaces them or marks them as discarded. TRIM closes that information gap. The operating system can notify the storage device that selected logical blocks no longer need their old contents. The SSD may then treat the associated data as disposable during its internal space-management work.

Database 15 Sep 2026 4 min read

PostgreSQL TOAST Moves Large Values Outside Heap Rows

PostgreSQL TOAST Moves Large Values Outside Heap Rows PostgreSQL heap tuples cannot span data pages. A row containing a large text, bytea, jsonb, or other variable-length value therefore cannot simply continue onto the next heap page. TOAST provides the storage mechanism that keeps such rows representable: eligible values can be compressed, moved out of line, or both. The visible SQL value does not change when this happens. The physical representation does. A heap tuple may contain a compact reference while the value’s bytes reside as chunk rows in a separate relation associated with the table.

Tech 15 Sep 2026 7 min read

NVMe Queues Let Storage Handle Many Commands in Parallel

NVMe storage does not send every read or write through one shared command line. The protocol is built around queue pairs: software places commands into a submission queue, and the controller reports finished work through a corresponding completion queue. That structure matters most when several processor cores and application threads are generating storage work at the same time. Multiple queues can distribute command handling across cores, reduce contention around a single software path, and keep a fast solid-state drive supplied with enough outstanding work.

Tech 15 Sep 2026 6 min read

Flash Wear Leveling Spreads Erase Cycles Across NAND Cells

NAND flash has a finite endurance budget. Programming changes stored charge, but cells cannot simply overwrite arbitrary existing data in place. Pages are programmed inside larger erase blocks, and an erase operation resets a block before its pages can accept fresh programming. That geometry creates an endurance problem. A workload may update the same logical address thousands of times even though the drive contains many other blocks that receive almost no writes. If every logical address stayed permanently tied to one physical location, a small hot region could reach its erase-cycle limit while much of the flash remained lightly used.

Tech 14 Sep 2026 5 min read

SSD TRIM Marks Unused Blocks Before New Writes

Deleting a file changes filesystem metadata, but a solid-state drive cannot infer from that metadata alone which stored pages no longer contain useful data. Without an extra signal, the drive may continue treating those pages as valid even after the operating system has released their logical addresses. TRIM provides that signal. It lets the operating system tell the SSD that selected logical block addresses no longer hold data that must be preserved. The controller can then treat the corresponding flash pages as disposable during later maintenance and write preparation.

Tech 14 Sep 2026 6 min read

SSD TRIM Marks Deleted Blocks for Reuse

Deleting a file and erasing data from flash memory are separate events on a solid-state drive. A file system can mark storage as free without immediately forcing the SSD to erase the corresponding flash cells. TRIM connects those two layers. It lets the operating system tell the drive that selected logical block addresses no longer contain data the system needs to preserve. File deletion changes the host view first A file system keeps metadata that maps files to logical storage locations. When a file is deleted, the file system can release those locations for later use. From the operating system’s perspective, the space becomes available quickly.

Tech 14 Sep 2026 4 min read

NVMe APST Moves Idle SSDs Into Lower Power States

An NVMe SSD can support several power states rather than operating at one fixed power level. Active states favor quick access and throughput, while deeper idle states can reduce energy use at the cost of extra time needed to return to full activity. Autonomous Power State Transition, commonly shortened to APST, lets the host configure automatic movement into selected lower-power states after defined idle periods. Once configured, the controller can perform those transitions without a separate host command for every idle event.