Skip to content

Archive

Distributed Systems

108 articles
Software Engineering 20 Sep 2026 7 min read

Deadline Propagation Stops Work After Callers Give Up

Deadline Propagation Stops Work After Callers Give Up A timeout at the edge does not automatically stop work deeper in a system. A client may abandon a request after two seconds while an API server continues waiting on another service, which may still be running a database query. The response has lost its consumer, yet CPU time, connections, memory, queue positions, and downstream capacity can remain occupied. Deadline propagation carries the caller’s time budget across those boundaries. Each component receives an absolute deadline or an equivalent remaining budget, refuses work that cannot start in time, and cancels operations when the budget expires. The goal is not merely faster failure. It is to keep useless work from surviving longer than the request that justified it.

Software Engineering 20 Sep 2026 6 min read

Deadline Propagation Stops Expired Requests from Consuming Downstream Capacity

A timeout placed only at the outer edge of a request does not automatically limit the work started deeper in the call graph. The client may stop waiting after 800 milliseconds while an internal service continues a database query, a remote call, or a queued task for several more seconds. The response is already useless to that client, yet the system is still spending capacity on it. Deadline propagation carries the request’s time boundary with the work. Each component can compare that boundary with its current clock, reserve time for its own processing, and refuse or cancel work that no longer fits. The result is not merely faster failure. It is a tighter relationship between useful work and resource consumption.

Software Engineering 20 Sep 2026 6 min read

Consistent Hashing Limits Key Movement When Nodes Change

Partitioning by hash(key) % N is simple when the node count stays fixed. The arithmetic becomes disruptive when N changes. Moving from four nodes to five changes the divisor, so many keys select a different remainder even though only one node joined. Consistent hashing changes the mapping. Keys and nodes are placed in the same circular hash space. A key belongs to the first node encountered in a chosen direction around the ring. Adding or removing a node changes ownership only for ranges adjacent to that membership change.

Software Engineering 20 Sep 2026 4 min read

Circuit Breakers Limit Repeated Calls to Failing Dependencies

Circuit Breakers Limit Repeated Calls to Failing Dependencies A remote dependency can fail in a way that is both slow and expensive. Requests wait for timeouts, workers remain occupied, retries add more traffic, and a local service can lose capacity even when its own code is healthy. A circuit breaker places a stateful decision in front of that call path. While the dependency behaves acceptably, calls pass through. After the configured failure condition is reached, the breaker opens and rejects new calls locally for a bounded period. Later, it admits a small number of probes before deciding whether normal traffic can resume.

Software Engineering 20 Sep 2026 5 min read

Bulkheads Isolate Concurrency Across Dependencies

Bulkheads Isolate Concurrency Across Dependencies A service can have plenty of CPU and still become unavailable because one dependency stops completing work. Requests waiting on a slow database, remote API, or storage service retain execution slots, connections, memory, and queue positions. If unrelated operations share the same finite pool, one saturated path can consume capacity needed by healthy paths. Bulkhead isolation divides that shared concurrency into explicit budgets. Calls to one dependency or workload class use a bounded pool that other classes cannot exhaust. The pattern does not repair a failing dependency. It limits the amount of local capacity that failure can occupy.

Software Engineering 20 Sep 2026 7 min read

Adaptive Concurrency Limits Follow Service Capacity

Adaptive Concurrency Limits Follow Service Capacity A service can become slower before it becomes unavailable. As in-flight work rises, CPU queues grow, connection pools fill, lock contention increases, and downstream calls accumulate. A fixed concurrency ceiling can protect the service, but one number rarely fits every operating condition. Capacity shifts with request mix, cache hit rate, dependency latency, deployment shape, and resource pressure. Adaptive concurrency control treats the admission limit as a value that can move. The controller observes recent service behavior, raises the limit while additional concurrency remains productive, and reduces it when latency indicates growing queues or saturation. The goal is not maximum concurrency. It is enough parallel work to use available capacity without allowing queues to dominate response time.

Software Engineering 19 Sep 2026 8 min read

Transactional Outbox Couples State Change to Message Intent

A service that updates its database and publishes an event to a message broker crosses two independent commit boundaries. If the database commit succeeds and the broker publish fails, durable state exists without its corresponding message. Reversing the order only reverses the failure: a consumer can observe a message for a state change that never commits. A transactional outbox narrows this gap by placing the application write and a durable message record in the same local database transaction. Publication moves to a separate relay. The pattern does not make the database and broker one atomic system; it changes the boundary so message intent becomes part of the database commit.

Software Engineering 19 Sep 2026 6 min read

Request Coalescing Collapses Concurrent Cache Misses into One Fill

A cache can reduce steady-state backend traffic yet amplify work at the instant a popular entry expires. If one hundred requests observe the same missing key before any replacement value is stored, a conventional lookup path can send one hundred equivalent reads to the origin. The cache is functioning according to its lookup rules; the amplification comes from concurrency around the empty interval. Request coalescing changes that interval. The first caller for a key starts the fill, while later callers for the same key attach to that in-flight operation instead of starting equivalent work. When the operation completes, its result is distributed to the waiting callers and, when appropriate, stored in the cache.

Software Engineering 19 Sep 2026 8 min read

MQTT and QUIC Solve Different Parts of a Chat Transport

MQTT and QUIC Solve Different Parts of a Chat Transport MQTT and QUIC are often placed in the same comparison table when discussing real-time chat. That comparison is convenient, but it collapses two different protocol layers into one choice. MQTT is an application-layer messaging protocol. It defines concepts such as clients, brokers, topics, subscriptions, retained messages, session state, and delivery quality of service. QUIC is a secure transport protocol over UDP. It provides connections, streams, flow control, loss recovery, encryption, and connection migration mechanisms.

Software Engineering 19 Sep 2026 6 min read

Idempotency Keys Bound Retry Safety to a Request Identity

A client can send the same logical operation more than once even when it intended one effect. A timeout after POST /payments leaves an ambiguous boundary: the server may have committed the payment while the client received no response. Retrying restores delivery, but an ordinary retry can create a second payment. An idempotency key changes the interface by giving repeated attempts a stable request identity. The key is not a substitute for transactionality, and it does not make every operation intrinsically idempotent. It creates a protocol between client and server: attempts carrying the same key are treated as candidates for the same logical operation. The server still needs rules for request equivalence, concurrent arrival, persistence lifetime, failure recovery, and response replay.

Software Engineering 19 Sep 2026 7 min read

Idempotency Keys Bind Retries to One Logical Mutation

A client can lose an HTTP response after the server has committed the requested mutation. From the client’s perspective, the operation is unresolved: the connection failed, but that failure does not reveal whether durable state changed. Retrying the same POST can then create a second order, payment attempt, reservation, or other mutation. An idempotency key gives the retry a stable identity that is separate from any single transport attempt. The server can associate repeated requests carrying that identity with one logical operation. That mechanism narrows an ambiguity at the API boundary, but the key alone is not a guarantee. Its scope, persistence, request comparison, concurrency control, and replay policy determine what repeated delivery actually means.

Software Engineering 19 Sep 2026 6 min read

Fencing Tokens Reject Stale Writers After Lease Expiry

A distributed lease can transfer ownership without stopping the process that previously held it. A worker may pause long enough for its lease to expire, then resume after another worker has acquired the same lease. At that point both processes can execute code that was written under the assumption of exclusive ownership. Lease expiry settles ownership in the coordination service. It does not revoke CPU time, cancel an in-flight network request, or erase buffered I/O on the former holder. Fencing tokens address that gap by carrying an ordering value from the ownership decision to the resource being protected.

Software Engineering 19 Sep 2026 6 min read

Fencing Tokens Close the Stale Lease Writer Gap

A distributed lease can expire while its holder is unable to run. The holder may later resume with local state that still says it owns the lease, even though another client has already acquired a newer lease. If the protected storage or service accepts operations solely because the client once acquired the lease, two clients can mutate the same resource across different points in time. A fencing token moves the decisive check from lease ownership into the protected resource. Each successful acquisition receives a token ordered after every earlier token. The resource records the greatest accepted token and rejects operations carrying an older value. The lease still coordinates acquisition, but the token constrains what a delayed former holder can do after it resumes.

Software Engineering 19 Sep 2026 7 min read

Deadline Propagation Bounds Request Lifetime Across Service Calls

A request can stop being useful before every process handling it stops working. An HTTP client may give up after two seconds while an upstream service continues a database query, an RPC, and a retry sequence for several more seconds. Those operations still consume connections, CPU time, queue capacity, and downstream concurrency even though their result no longer has a recipient. A deadline makes that usefulness boundary explicit. Propagating it through nested calls gives participating components a common upper bound derived from the original request. This differs from assigning an independent timeout at every hop: local timeouts limit individual operations, while a propagated deadline limits the lifetime of the operation graph.

Software Engineering 19 Sep 2026 7 min read

Circuit Breakers Bound Failure Traffic Across Service Calls

A circuit breaker changes the admission decision for an outbound call before the dependency receives it. In the closed state, calls proceed and their outcomes feed a failure policy. Once that policy trips, the breaker enters the open state and rejects subsequent calls locally. After a configured recovery interval, a limited set of probe calls can test whether the dependency is usable again. That mechanism is distinct from retries. A retry issues another attempt after a failed attempt. A breaker can prevent an attempt from being issued at all. Combining the two without a precise ordering can amplify traffic during an outage or keep a breaker open based on signals that do not represent dependency health.

Software Engineering 19 Sep 2026 6 min read

Circuit Breakers Bound Failure Amplification Across Service Calls

A service call can fail quickly and still create a larger system problem. When every upstream request continues to invoke a downstream dependency that is already failing, each attempt consumes connection capacity, worker time, retry budget, and queue space. The dependency receives traffic it cannot currently serve, while callers spend resources waiting for outcomes that are already strongly correlated with recent failures. A circuit breaker puts a stateful decision boundary in front of that call. Instead of treating every request as an independent opportunity to try the dependency, it records recent failure state and can reject calls locally for a bounded interval. Recovery is then tested through controlled probes rather than a full return of traffic.

Software Engineering 16 Sep 2026 7 min read

HTTP Stale-While-Revalidate Moves Cache Refresh Off the Request Path

A cache can return an expired stored response immediately and start validation in parallel when stale-while-revalidate permits that reuse. The request that encounters the stale entry therefore does not have to inherit origin validation latency, but it can receive representation data older than the normal freshness lifetime. This is a deliberate shift in the cache contract. Freshness still expires at the configured boundary. The extension adds a separate interval in which stale reuse is permitted while validation proceeds, so response age and request latency become partially decoupled.

Software Engineering 16 Sep 2026 8 min read

HTTP If-Range Couples Partial Retrieval to Representation Identity

A client that has only part of an HTTP representation faces a consistency problem when it asks for the missing bytes later. Byte offsets are meaningful only against the representation whose bytes established those offsets. If the selected representation changes between requests, combining an old prefix with a new suffix can produce data that no server ever emitted. If-Range attaches representation identity to that partial-retrieval boundary. When its validator matches, the server can process the accompanying Range field. When it does not match, the server ignores Range and sends the complete selected representation through the normal successful response path instead of returning a failed-precondition response.

Software Engineering 16 Sep 2026 8 min read

HTTP If-Match Turns Representation State Into a Write Precondition

An HTTP origin can refuse a PUT or DELETE before applying it when the request carries If-Match and the selected representation no longer has an accepted entity tag. The condition converts a representation validator into a write precondition: a client can say that a mutation is valid only against state matching a version it previously observed. This mechanism addresses a specific concurrency boundary. It can prevent one client from silently replacing resource state after another client has changed the selected representation. It does not turn HTTP into a transaction protocol, lock the resource between requests, or guarantee that an entity tag represents every piece of application state involved in a mutation.

Software Engineering 16 Sep 2026 7 min read

HTTP 421 Misdirected Request Marks a Connection Authority Boundary

An HTTP/2 client can reuse one secured connection for requests to more than one origin when the server is authoritative for those origins. A request can still reach a server instance whose connection context does not fit the target URI. 421 Misdirected Request exists for that boundary: the server rejects the routing context rather than treating the target resource itself as missing. This distinction separates resource semantics from connection authority. A 421 response says that this server, on this path or connection context, is unable or unwilling to produce an authoritative response for the target URI. It does not say that the resource has been deleted, that its method is forbidden, or that the request representation is invalid.

Software Engineering 15 Sep 2026 6 min read

Tombstones Preserve Deletions Across Replica Gaps

A replicated store cannot always represent deletion as immediate absence. If one replica removes a record while another replica is disconnected, erasing every trace of the record also erases the evidence needed to distinguish a deliberate deletion from a replica that simply has not seen recent state. A tombstone keeps that evidence as versioned metadata. Instead of removing the key from the replication domain at once, the system records a deletion marker that participates in reconciliation. A replica carrying an older live value can then compare its state with the marker and discard the obsolete value.

Software Engineering 15 Sep 2026 7 min read

Request Coalescing Turns Concurrent Cache Misses Into Shared Work

A cache entry can expire while hundreds of requests for the same key are already in flight. If every request observes the miss independently, each can start the same backend operation before any result reaches the cache. The cache still limits work across time, but it does not limit duplicate work during that miss interval. Request coalescing adds a second boundary: concurrent operations for the same logical key can share one in-flight computation. One caller becomes the active producer, while matching callers wait for that producer’s result instead of starting equivalent work. The mechanism is also called single-flight suppression in systems that expose it as a concurrency primitive.

Software Engineering 15 Sep 2026 7 min read

Idempotency Keys Bind Retries to One Logical Operation

A client can transmit a state-changing request, lose the response, and retry without knowing whether the first attempt committed. At that boundary, transport failure has created ambiguity rather than proof of application failure. Repeating the mutation blindly can create a second logical effect. An idempotency key gives the server a stable operation identity across those delivery attempts. The first accepted request associates the key with an operation record. A later request carrying the same key can then reuse the recorded outcome instead of executing the mutation again.

Software Engineering 15 Sep 2026 7 min read

Fencing Tokens Reject Stale Lock Holders

A process acquires a distributed lease, pauses long enough for that lease to expire, then resumes. Another process has already acquired the same lease. At that moment both processes can execute code that was entered under an apparently valid acquisition, even though only the newer owner should retain authority. The lease itself cannot retract instructions from the paused process. Expiration changes coordination state; it does not erase local state, stop a suspended runtime, or cancel an operation already queued elsewhere. This gap is the central limitation of treating a distributed lock as a remote version of an in-process mutex.