Skip to content

Archive

Software Engineering

462 articles
Software Engineering 15 Sep 2026 6 min read

Atomic Rename Separates Visibility From Crash Durability

A successful rename() can replace an existing pathname without exposing an interval in which that destination name is absent. That visibility property makes rename a common publication boundary for configuration files, checkpoints, manifests, and other file-backed state. It does not, by itself, establish that the replacement survives an abrupt loss of power. The distinction is between namespace atomicity and persistence. Atomic replacement constrains what concurrent observers can see while the system is running. Crash durability concerns which writes and metadata changes are guaranteed to remain after volatile state disappears. Treating those properties as equivalent creates a gap precisely at the failure boundary that atomic replacement is often intended to protect.

Software Engineering 14 Sep 2026 7 min read

Transactional Outboxes Move Atomicity Into the Database

A service updates an order row and emits an event about that update. If the database commit succeeds but the broker publish fails, durable state says one thing while downstream consumers receive no corresponding message. Reversing the order only moves the gap: a successful publish followed by a failed database transaction exposes an event for state that never committed. The difficulty is not message syntax or retry configuration. It is atomicity across two systems that do not share a transaction. A transactional outbox changes the boundary. The application writes its domain state and a message record into the same database transaction, then a separate relay publishes committed outbox records to the broker.

Software Engineering 14 Sep 2026 7 min read

Poison Messages Turn Retries Into Queue Retention

A queue consumer receives a message, rejects it, and receives the same message again. That cycle is useful when the rejection came from a transient condition. It is structurally different when the payload can never be processed by the current consumer. The broker can keep honoring redelivery semantics while the application makes no forward progress on that message. Such a message is commonly called a poison message. The important property is not that it contains malformed bytes. A syntactically valid message can be permanently unprocessable because its schema is unsupported, a required invariant is violated, referenced data can never exist, or application logic deterministically rejects its state.

Software Engineering 14 Sep 2026 9 min read

Deadlines Shrink Across Service Boundaries

A service receives a request with 480 milliseconds remaining before its deadline. It spends 90 milliseconds reading state, then calls another service with a fixed 500-millisecond timeout. The downstream call can now outlive the request that caused it. Nothing about either timeout is internally inconsistent; the inconsistency appears at the boundary between them. Timeouts are often configured as local limits: a database query gets one value, an HTTP client another, a queue operation a third. A deadline represents a different constraint. It gives an operation an end point, so every later stage can compare its own work against the same finite lifetime.

Software Engineering 13 Sep 2026 9 min read

Write Skew Across Disjoint Rows

Write Skew Across Disjoint Rows Two transactions read the same set of rows, reach compatible decisions, and then update different rows. Neither transaction overwrites the other’s write. Both commits can still leave the database in a state that violates a rule spanning those rows. That shape is write skew. It is easy to miss because many concurrency discussions center on two writers contending for one row. Write skew has no such collision. The conflict exists at the level of an invariant inferred from several records, while the physical writes remain disjoint.

Software Engineering 13 Sep 2026 10 min read

Version Columns Turn Lost Updates Into Conflicts

Two transactions can read the same row, compute different changes, and then write in sequence. If each update replaces values derived from its earlier read, the later write can erase part of the earlier one without either transaction observing a database error. A version column changes that interaction. The row carries a generation value alongside its domain fields, and an update is accepted only when the generation still matches the value observed by the writer. A stale writer no longer looks identical to a current writer at the storage boundary.

Software Engineering 13 Sep 2026 7 min read

Schema Changes Are Multi-Version Protocols

A column rename looks atomic in a schema diff. A deployed system rarely experiences it that way. During a rolling release, old application processes can remain active after new processes start. Background jobs may run code built from another release. Replicas can lag behind a primary. Queued work can outlive the binary that created it. Data written before the change remains present after the new schema exists. The migration therefore crosses several versions of code and data at once.

Software Engineering 13 Sep 2026 9 min read

Request Coalescing Turns Cache Miss Bursts Into Shared Work

A cache entry expires at a single instant, but requests for that entry do not necessarily arrive one at a time. If twenty callers observe the same miss before any caller has repopulated the cache, a conventional cache-aside path can send twenty equivalent reads to the origin. The cache is functioning according to its contract; the concurrency around the miss is creating duplicated work. Request coalescing changes that boundary. Instead of treating each miss as permission to start an origin operation, callers for the same key can share one in-flight operation. One caller becomes the producer of the pending result. Other callers wait for that result rather than starting equivalent work.

Software Engineering 13 Sep 2026 7 min read

Monotonic Time Belongs in Elapsed-Time Measurement

A timeout can be represented by an ordinary timestamp comparison: record a start value, read the clock later, subtract, and compare the result with a limit. The arithmetic looks complete. The clock model is not. Civil time exists to place events on a shared calendar. It can be corrected to stay aligned with an external time reference. Elapsed-time measurement has a different requirement: later observations within one running system need an ordering suitable for measuring an interval. A clock correction that improves civil-time accuracy can therefore be harmful when the same reading is treated as a stopwatch.

Software Engineering 13 Sep 2026 7 min read

Keyset Pagination Under Concurrent Writes

Keyset Pagination Under Concurrent Writes A query returns twenty rows ordered by creation time. Before the client asks for the next twenty, another transaction inserts a row near the front of that order. The data set has changed, but the client still expects page two to continue from the point page one reached. That expectation exposes the main difference between offset pagination and keyset pagination. An offset identifies a position in a particular query result. A keyset cursor identifies an ordering boundary. Under concurrent writes, those are not equivalent references.

Software Engineering 13 Sep 2026 7 min read

Idempotent Consumers and Durable Duplicate Detection

Idempotent Consumers and Durable Duplicate Detection A consumer commits a database transaction, then loses its connection before acknowledging the message. The broker has no evidence that processing finished, so a delivery protocol that permits redelivery can present the same message again. The second delivery is not evidence that the first transaction failed. From the consumer’s perspective, the important fact is more precise: message delivery and application commit have separate completion points. If the broker cannot atomically participate in the application’s state transition, an acknowledgement can be lost after the application effect is already durable.

Software Engineering 13 Sep 2026 7 min read

Garbage Collection Does Not Close External Resources

An object can become unreachable while a file descriptor, socket, database transaction, or operating-system lock associated with it still has a meaningful external lifetime. Garbage collection can reclaim managed memory after reachability disappears. It does not, by that fact alone, perform the protocol operation that releases an external resource. The distinction is structural. A collector reasons about references inside a managed heap. A resource such as a file descriptor is an entry maintained by the operating system, and a database transaction is state maintained by another component. Their release semantics come from APIs and protocols outside the collector’s reachability model.

Software Engineering 13 Sep 2026 6 min read

Fencing Tokens Make Expired Leases Observable

A process can hold a distributed lease, pause long enough for that lease to expire, then resume with local state that still says it owns the resource. Another process may already have acquired a newer lease during the pause. At that point, mutual exclusion in the lock service is not enough: two processes can each act as if they have authority, even though only one lease is current. This is a boundary problem between coordination and the resource being protected. A lease service can decide which holder is current according to its own state. It cannot retroactively erase instructions already held by an old process, nor can it stop that process from sending a request after a long pause.

Software Engineering 13 Sep 2026 8 min read

Circuit Breakers Bound Failed Call Admission

A remote call can fail in a few milliseconds or consume its entire timeout budget before returning an error. If callers keep issuing equivalent requests while the dependency remains unable to serve them, each attempt spends resources on an outcome that recent evidence already suggests is unavailable. Retries can increase that pressure because one logical operation may create several physical calls. A circuit breaker changes call admission rather than the remote protocol. It records recent outcomes, moves between explicit states, and can reject new calls locally for a bounded period. After that period, it permits limited probes to test whether normal traffic can resume.

Software Engineering 13 Sep 2026 7 min read

Canonical Serialization Makes Byte Identity Explicit

Two serialized documents can represent the same application value and still differ byte for byte. An object member can appear in another order. A number can use a different textual form. Unicode text can contain distinct code-point sequences that render alike. Whitespace may be optional. A serializer can make any of these choices while remaining valid for its format. That flexibility is usually harmless when serialization is only a transport boundary. It becomes part of system semantics when bytes are hashed, signed, compared, cached by digest, or used as content addresses. At that point, logical equivalence is not enough. The operation consumes an exact byte sequence.

Software Engineering 13 Sep 2026 9 min read

Cancellation Is a Protocol, Not a Thread Kill

A caller can stop waiting for a result while the operation producing that result continues to run. The distinction is easy to miss because many APIs expose cancellation through a single method, token, context, or signal. That surface can look like a command to terminate work. In most cooperative designs, it is closer to a state transition: further work is no longer wanted. The gap matters once an operation owns resources, crosses process boundaries, or has already produced side effects. A cancelled HTTP request does not retroactively erase a committed database transaction. A task that notices a cancellation token between two writes cannot make the first write disappear. A parent that abandons a child computation also needs a rule for who observes the child’s eventual completion and who releases anything the child owns.

Software Engineering 12 Sep 2026 8 min read

Write-Ahead Logging and the Meaning of Commit

A database can report a transaction as committed while the data pages touched by that transaction are still absent from their final locations on disk. That behavior is not a contradiction. In systems built around write-ahead logging, durability is established by the log before the modified pages need to reach durable storage. The distinction matters because a transaction changes several kinds of state at once. It changes the logical database, it changes in-memory page images, and it creates recovery information. Treating those as a single physical write obscures the mechanism that gives commit its meaning after a crash.

Software Engineering 12 Sep 2026 8 min read

Vector Clocks and the Shape of Concurrent State

A single version number can state that one value came after another only when every update participates in the same ordered sequence. Replicated state breaks that assumption as soon as independent writers can accept changes without first agreeing on one global next version. Two replicas can each move forward from the same ancestor. Calling one state version 8 and the other version 9 creates an order, but that order may describe the numbering scheme rather than the causal relation between the writes. Vector clocks represent a different fact: which update history a state has observed.

Software Engineering 12 Sep 2026 7 min read

Transactional Outbox: Moving Publication Across the Commit Boundary

Transactional Outbox: Moving Publication Across the Commit Boundary A service that changes database state and publishes a message has two distinct side effects. A local transaction can make the database change atomic, and a broker can accept the message durably, but those facts do not make the pair atomic. The awkward interval sits between them. If the database commit succeeds and publication does not, durable state exists without the corresponding message. Reversing the order only reverses the exposure: a message can become visible before the database commit succeeds.

Software Engineering 12 Sep 2026 9 min read

Transactional Outbox at the Database-Broker Boundary

Transactional Outbox at the Database-Broker Boundary A database commit and a broker publish are two separate state transitions. An application can complete either one first, but unless both systems participate in a common transaction protocol, there is an interval in which one side has changed and the other has not. That interval is the central problem behind the transactional outbox pattern. The pattern does not make a database and broker commit atomically. Instead, it moves the durable publication decision into the same database transaction as the application state change. A separate publisher later converts that recorded intent into a broker message.

Software Engineering 12 Sep 2026 8 min read

The ABA Problem Behind Successful Compare-and-Swap

A compare-and-swap operation answers a narrow question: does a memory location contain the expected bit pattern at the instant of the atomic operation? If it does, the replacement can proceed. The operation does not establish that the location remained unchanged between an earlier read and the later comparison. That distinction creates the ABA problem. A thread observes value A, pauses, and later performs a compare-and-swap expecting A. During the pause, other work changes the location from A to B and then back to A. The comparison succeeds because the current representation matches the expected representation, even though the shared state passed through a transition that may invalidate assumptions attached to the first observation.

Software Engineering 12 Sep 2026 12 min read

Saga Transactions: Coordinate Multi-Service Changes with Compensation

Saga Transactions: Coordinate Multi-Service Changes with Compensation A business operation can cross several services even when no single database transaction spans them all. An order flow might reserve inventory, authorize payment, create a shipment, and confirm the order. Each service owns its data and commits independently. That independence creates a difficult failure case. Inventory can be reserved successfully, then payment authorization can fail. A database rollback in the order service cannot undo a committed reservation in the inventory service.

Software Engineering 12 Sep 2026 8 min read

Retry Amplification and the Role of Jitter

Retry Amplification and the Role of Jitter A failed request can create more traffic than a successful one. If a caller immediately repeats an operation after a transient error, the original unit of demand becomes two attempts. Add another retrying layer above that caller, and a single logical request can fan out into several physical attempts before any component has recovered. Retries are often described as a way to tolerate temporary faults. That description is incomplete because retry behavior also changes load. The mechanism sits inside a feedback loop: failure triggers another attempt, another attempt consumes capacity, and consumed capacity can affect the conditions that produced the failure.

Software Engineering 12 Sep 2026 10 min read

Request Coalescing at Hot Cache Misses

Request Coalescing at Hot Cache Misses A cache entry expires at 22:00:00.000. Ten milliseconds later, two hundred requests ask for the same key. A cache-aside implementation sees two hundred misses. If every caller independently reads the backing service, a single expiration event becomes two hundred concurrent backend operations. Nothing is wrong with the cache lookup itself. The amplification comes from treating identical in-flight work as unrelated. Request coalescing changes that boundary. Callers that need the same absent key share one active load, while requests for other keys continue independently. The mechanism is small, but its semantics reach beyond a mutex: it defines which operations may share a result, how failures fan out, what cancellation means, and when another load may begin.