Skip to content

Archive

Distributed Systems

108 articles
Software Engineering 21 Sep 2026 6 min read

Hedged Requests Cut Tail Latency at a Capacity Cost

Hedged Requests Cut Tail Latency at a Capacity Cost A service can have acceptable median latency while a small fraction of requests take much longer. Queueing, a cold cache, runtime pauses, transient packet loss, or a slow storage operation can leave one attempt far behind the normal path. At sufficient fan-out, those rare delays become common at the aggregate request boundary. A hedged request starts a second equivalent attempt after the first has remained incomplete for a selected delay. The caller accepts the first useful result and cancels or discards the other attempt.

Software Engineering 21 Sep 2026 6 min read

Fencing Tokens Stop Stale Lease Holders from Writing

Fencing Tokens Stop Stale Lease Holders from Writing A distributed lease can decide which client currently owns a resource, but lease expiry does not instantly stop the previous holder. A process can pause, lose network access, or stall long enough for its lease to expire. Another client then acquires the lease. If the old process resumes and still has access to the protected storage or service, both clients can issue writes.

Software Engineering 21 Sep 2026 5 min read

Deadline Propagation Preserves Timeout Budgets Across RPC Hops

Deadline Propagation Preserves Timeout Budgets Across RPC Hops A timeout that restarts at every service boundary can turn a short caller budget into a much longer chain of work. A client may allow 800 milliseconds, service A may spend 500 milliseconds locally, then call service B with a fresh 800-millisecond timeout. B can continue working long after the client has stopped waiting. Deadline propagation keeps one end time attached to the request. Each hop derives its remaining budget from that deadline and refuses to start work that cannot fit within it. The result is not faster execution by itself. It is bounded execution that respects the time contract established upstream.

Software Engineering 21 Sep 2026 6 min read

Deadline Propagation Preserves Request Time Budgets

Deadline Propagation Preserves Request Time Budgets A request can cross several services before producing a response. Each hop may have its own queue, network call, retry policy, and local timeout. If those limits are chosen independently, the total path can run far longer than the caller is prepared to wait. Deadline propagation gives the path one temporal boundary. The initiating caller supplies an absolute deadline, or a time budget that is converted into one. Each downstream component uses the remaining interval rather than starting a fresh timeout from zero.

Software Engineering 21 Sep 2026 6 min read

Consistent Hashing Limits Key Movement During Membership Changes

Consistent Hashing Limits Key Movement During Membership Changes A simple hash partition often looks sufficient: owner = hash(key) % node_count With four nodes, every key maps to one of four remainders. The problem appears when the membership changes. Moving from four nodes to five changes the modulus, so a large fraction of keys select a different owner even though only one node was added.

Software Engineering 21 Sep 2026 6 min read

Circuit Breakers Limit Cascading Failure

Circuit Breakers Limit Cascading Failure A slow or failing dependency can consume more than its own capacity. Callers wait, retry, hold sockets, occupy worker slots, and retain memory while requests remain unresolved. As pressure spreads upstream, a local fault can become a service-wide saturation event. A circuit breaker places a stateful gate around calls to that dependency. It observes outcomes, opens when the configured failure policy is met, rejects calls for a period, then permits a small number of probes. Successful probes can return the breaker to normal traffic; failed probes send it back to the open state.

Software Engineering 21 Sep 2026 7 min read

Circuit Breakers Bound Calls to an Unhealthy Dependency

Circuit Breakers Bound Calls to an Unhealthy Dependency A remote dependency can fail in a way that is both persistent and expensive. Connections time out, worker slots remain occupied, request queues grow, and retries add more traffic to a service that is already unable to respond. A caller that keeps issuing the same class of request can turn one dependency failure into pressure across its own process. A circuit breaker places a stateful admission decision in front of those calls. While the dependency behaves within policy, calls pass through. After enough qualifying failures, the breaker opens and rejects new calls locally for a period. Recovery is tested with limited traffic rather than a full return to normal load.

Software Engineering 21 Sep 2026 7 min read

Bulkheads Isolate Concurrency Across Failure Domains

Bulkheads Isolate Concurrency Across Failure Domains A service can remain reachable while its useful capacity disappears. A slow dependency holds requests open, those requests occupy workers or connection slots, and unrelated traffic waits behind work that cannot finish promptly. The fault began in one path, but a shared resource pool lets it consume capacity needed by every path. Bulkhead isolation partitions that finite capacity. Calls associated with one failure domain receive a bounded share rather than competing without separation for the entire pool. When one partition fills, admission fails or waits within that partition while capacity assigned to other work remains available.

Software Engineering 21 Sep 2026 6 min read

Backpressure Bounds Work When Consumers Fall Behind

Backpressure Bounds Work When Consumers Fall Behind A fast producer and a slower consumer can coexist for a short burst if a buffer absorbs the difference. The same arrangement becomes unstable when the rate mismatch persists. Pending work accumulates, memory rises, latency stretches, and items may expire before a consumer reaches them. Backpressure changes the contract between both sides. Instead of accepting work without regard to downstream state, the system exposes limited capacity to the producer. When that capacity is exhausted, production pauses, admission is rejected, or another explicit overload policy takes effect.

Software Engineering 21 Sep 2026 7 min read

Adaptive Concurrency Limits Track Available Service Capacity

Adaptive Concurrency Limits Track Available Service Capacity A fixed concurrency ceiling is easy to operate when service capacity is stable. Real systems rarely stay in one operating regime. Database contention, cache hit rate, request mix, downstream latency, CPU availability, and deployment changes can all move the amount of work a service can sustain at once. Adaptive concurrency control treats the in-flight limit as a control variable. The limiter admits work up to a current ceiling, observes service behavior, then adjusts that ceiling. The aim is not maximum concurrency. It is enough concurrency to use available capacity without allowing queues to grow far beyond the useful operating region.

Software Engineering 20 Sep 2026 6 min read

Visibility Timeouts Turn Message Delivery into a Renewable Lease

A queue consumer often needs time to perform work before it can safely acknowledge a message. Removing the message at receive time would make a consumer crash capable of losing work. Keeping it immediately available would let several consumers process the same item at once. A visibility timeout occupies the middle ground. Receiving a message makes it temporarily unavailable to competing consumers. The consumer gets a bounded interval to finish and acknowledge it. If that interval expires first, the queue can expose the message for another delivery.

Software Engineering 20 Sep 2026 6 min read

Transactional Outbox Keeps Database State and Events Aligned

A service often needs one operation to change database state and emit an event. An order may move to paid while OrderPaid must reach a message broker. Those two writes cross different systems, so a normal database transaction cannot make both commits atomic. Writing the database first leaves a gap: the process can stop after commit but before publishing. Publishing first creates the opposite gap: consumers can observe an event for a database change that later fails.

Software Engineering 20 Sep 2026 7 min read

Tombstones Preserve Deletes Across Replicas Until Safe Garbage Collection

Tombstones Preserve Deletes Across Replicas Until Safe Garbage Collection Deleting a value from one copy of replicated data is not enough to delete it from the system. Another replica may be offline, delayed, or partitioned when the delete occurs. If the active replica simply removes the record, it also removes the evidence that a deletion happened. A stale replica can later return with an older value and make that value visible again.

Software Engineering 20 Sep 2026 6 min read

Token Buckets Separate Sustained Rate from Burst Capacity

A rate limit expressed only as “100 requests per second” leaves an important policy question open. Can a client send 100 requests at the first instant of each second, or must those requests be spread evenly? A token bucket makes that distinction explicit by separating sustained rate from burst capacity. The limiter maintains a balance of tokens up to a fixed capacity. Tokens arrive at a configured refill rate. An operation is admitted only when enough tokens are available, and admission deducts its cost from the balance. Idle time accumulates capacity for a later burst, but never beyond the bucket limit.

Software Engineering 20 Sep 2026 6 min read

Stale-While-Revalidate Keeps Cache Refresh off the Request Path

Stale-While-Revalidate Keeps Cache Refresh off the Request Path A cache entry does not become useless at the exact instant its freshness timer expires. For some data, a value that is a few seconds old is still preferable to making every caller wait for a backend refresh. Stale-while-revalidate uses that tolerance explicitly: the cache may serve an expired value for a bounded interval while a refresh runs separately. The policy changes refresh from a request-path requirement into background work for entries that remain acceptable while stale. It can reduce latency spikes around expiration, but only when the application can state how stale a response may become.

Software Engineering 20 Sep 2026 4 min read

Request Coalescing Collapses Concurrent Cache Misses

Request Coalescing Collapses Concurrent Cache Misses A cache miss can become expensive when many requests ask for the same key at nearly the same time. Without coordination, each caller can start an identical database query, remote call, or computation. The cache eventually fills, but the backend absorbs a burst precisely when the cached value is absent. Request coalescing changes the concurrency boundary. The first caller starts the load and publishes an in-flight entry for that key. Later callers join that entry instead of starting equivalent work. When the load finishes, its result is distributed to the waiting callers and the in-flight entry is removed.

Software Engineering 20 Sep 2026 6 min read

Rendezvous Hashing Keeps Key Placement Stable as Nodes Change

Rendezvous Hashing Keeps Key Placement Stable as Nodes Change Distributed systems often need a deterministic answer to a placement question: given a key and a current set of nodes, which node owns the key? A simple modulo rule such as hash(key) % N is compact, but changing N can move a large fraction of keys at once. Rendezvous hashing, also called highest-random-weight hashing, uses a different rule. For each key, it computes a deterministic score for every eligible node and selects the node with the highest score. Adding or removing a node changes placement only for keys whose ranking is affected by that membership change.

Software Engineering 20 Sep 2026 5 min read

Power of Two Choices Reduces Load Imbalance with Two Samples

Power of Two Choices Reduces Load Imbalance with Two Samples A load balancer that chooses one destination uniformly at random is cheap and decentralized, but random placement can produce uneven queues. At the other extreme, selecting the least loaded destination from the entire pool requires current load information for every candidate and can make the balancer itself expensive. The power-of-two-choices strategy sits between those designs. For each request, sample two eligible destinations, compare a load signal, and send the request to the better candidate. Two observations are enough to avoid many unlucky placements without requiring a global search.

Software Engineering 20 Sep 2026 6 min read

Lease Renewal Needs a Safety Margin Before Expiry

Lease Renewal Needs a Safety Margin Before Expiry A lease grants a holder temporary authority until a recorded expiry. Keeping that authority requires renewal before the deadline. Scheduling renewal at the deadline itself leaves no room for network delay, scheduler pauses, storage latency, or a transient retry. A safer design attempts renewal earlier. The interval between the planned renewal and expiry is a safety margin: time reserved for ordinary uncertainty before the lease is treated as lost.

Software Engineering 20 Sep 2026 7 min read

Idempotency Keys Make Retried Writes Safe to Repeat

Idempotency Keys Make Retried Writes Safe to Repeat A client can lose the response to a successful write. The connection may close after the server commits a payment, creates an order, or schedules a job but before the response reaches the caller. From the client’s perspective, failure and success can look identical. Retrying blindly is dangerous for operations with non-idempotent effects. Sending the same POST twice may create two resources or charge twice. Refusing to retry leaves the caller with an ambiguous outcome.

Software Engineering 20 Sep 2026 8 min read

Idempotency Keys Make Retried Commands Safe

A client can lose the response to a command even when the server completed the operation. The connection may close after a payment is recorded, a job is created, or an order is accepted. From the client’s point of view, timeout does not reveal whether the command failed before execution or succeeded before the response disappeared. Retrying is necessary for availability, but repeating a state-changing command can duplicate the side effect. An idempotency key gives all attempts for one logical command the same identity. The server stores the outcome associated with that identity and reuses it when the same command arrives again.

Software Engineering 20 Sep 2026 5 min read

Hedged Requests Trade Extra Work for Lower Tail Latency

A service can have a healthy median latency and still produce occasional requests that take far longer than the rest. Queueing, a slow replica, connection setup, garbage collection, storage stalls, or transient network delay can leave one attempt behind while equivalent capacity elsewhere remains available. A hedged request limits exposure to that single slow path. The client starts one attempt normally. If it is still pending after a configured delay, the client may start a second equivalent attempt. The first acceptable result is used, and the remaining attempt is cancelled when cancellation is supported.

Software Engineering 20 Sep 2026 6 min read

Hedged Requests Cut Tail Latency at a Controlled Cost

Hedged Requests Cut Tail Latency at a Controlled Cost A service can have acceptable median latency and still produce a small set of very slow responses. Queueing, runtime pauses, storage contention, packet loss, or a temporarily busy replica can push individual requests far beyond the common case. Hedged requests address that tail by sending a second copy after the first request has been outstanding for a chosen delay. The caller accepts the first valid response and cancels the remaining attempt. The technique trades a bounded amount of extra work for a chance to escape an unusually slow execution path.

Software Engineering 20 Sep 2026 7 min read

Fencing Tokens Block Stale Lock Holders at the Resource

Fencing Tokens Block Stale Lock Holders at the Resource A distributed lease can expire while its holder is still running. A process may pause for garbage collection, lose contact with the coordinator, stall under scheduler pressure, or resume after a machine suspension. The lock service can correctly grant the lease to another worker while the old worker still has unfinished work. That gap matters when both workers can reach the protected resource. A lease controls ownership in the coordinator; it does not automatically revoke a delayed database connection, storage request, or RPC that was prepared by the previous holder.