Skip to content

Archive

Distributed Systems

108 articles
Software Engineering 09 Sep 2026 11 min read

Using Hedged Requests to Reduce Tail Latency

Most calls to a dependency may finish quickly while a small fraction take much longer. A request that depends on one of those slow calls inherits the delay even when another healthy instance could have answered sooner. Increasing the timeout does not solve this problem. Retrying only after the timeout may also be too late: by then, the caller has already spent most of its latency budget. A hedged request is a deliberately delayed duplicate of an operation that is still in progress. The original request starts normally. If it has not completed after a chosen delay, the caller sends one additional equivalent request, usually to another eligible instance. The first acceptable result wins, and the remaining work is cancelled or ignored.

Software Engineering 09 Sep 2026 11 min read

Propagating Deadlines Through Call Chains

A request can have a timeout at every network call and still take far longer than the caller intended. The problem appears when each layer starts a fresh timeout. A frontend gives service A 800 milliseconds. Service A spends 300 milliseconds doing local work, then gives service B another 800 milliseconds. Service B spends 250 milliseconds and gives service C yet another 800 milliseconds. Every individual timeout looks reasonable, but the chain no longer has an 800-millisecond limit.

Python 08 Sep 2026 7 min read

Generate Time-Ordered IDs with Python UUIDv7

Random UUIDs are convenient identifiers: they can be generated without coordinating with a database, and the probability of collision is tiny. But a UUIDv4 primary key has one awkward property for ordered indexes: newly generated values are spread across the key space instead of tending toward the end of the index. UUID version 7 keeps the decentralized 128-bit UUID shape while putting a Unix-epoch millisecond timestamp at the front. Python 3.14 adds uuid.uuid7() to the standard library, so applications no longer need a third-party package just to generate RFC 9562 UUIDv7 values.

Software Engineering 08 Sep 2026 11 min read

Designing Read-Your-Writes Consistency

A user changes their delivery address, sees a success message, and opens the order page. The old address appears. They refresh a few seconds later and the new value finally shows up. Nothing necessarily lost the write. The system may have accepted it correctly and then served the next read from a copy that had not caught up yet. From the user’s perspective, however, a successful update appeared to reverse itself.

Software Engineering 07 Sep 2026 11 min read

Understanding Two-Phase Commit

Suppose one business operation must update two independent transactional resources. Writing to the first and then the second creates an uncomfortable failure case: the first write may commit while the second fails. Reversing the order only moves the problem. Two-phase commit, usually shortened to 2PC, is a coordination protocol for making one commit-or-abort decision across multiple participants that can each prepare and commit a local transaction. Its purpose is atomicity across those participants: under the protocol’s assumptions, they do not intentionally finish with some participants committed and others aborted for the same transaction.

Software Engineering 06 Sep 2026 11 min read

Design Retry Budgets to Prevent Retry Storms

Retries are one of the simplest reliability tools in distributed systems. A transient network failure, a short leader election, or a momentary overload can make a request fail even though the dependency becomes healthy again a few milliseconds later. Retrying can hide that temporary failure from the user. The same mechanism can also make an outage worse. If a struggling dependency starts failing requests and every caller retries immediately, the dependency receives extra work precisely when it has the least capacity to handle it. A small failure rate can turn into a retry storm: retries create more load, more load creates more failures, and those failures create still more retries.

Cloud Computing 04 Sep 2026 11 min read

Design Dead-Letter Queues for Poison Messages

Retries are useful when a failure is temporary. A database may be unavailable for a few seconds, a downstream service may return an overload response, or a network connection may disappear and recover. Retries become harmful when the message itself cannot succeed. A malformed payload, an unsupported schema version, a reference to permanently missing data, or a deterministic application bug can make the same message fail on every delivery. If the broker keeps returning that message to consumers indefinitely, the system spends capacity repeating work that has no chance of succeeding.

Software Engineering 02 Sep 2026 7 min read

Load Shedding and Bounded Queues for Overload Control

A service can be healthy at 500 requests per second and unusable at 700. The extra load does not merely make every request 40 percent slower. Queues grow, deadlines expire while work is still waiting, memory usage rises, retries create more traffic, and useful requests compete with work that can no longer finish in time. Overload control keeps that failure mode bounded. Instead of accepting unlimited work, a service limits concurrency and queueing, then rejects or degrades excess work early enough for the remaining requests to succeed.

Cloud Computing 01 Sep 2026 4 min read

Propagate Timeout Budgets Across Cloud Services

A request that crosses several cloud services does not have one timeout. It has a chain of deadlines: client, edge proxy, application, database, and downstream APIs. When those limits are configured independently, an upstream service can give up while downstream work continues consuming connections and CPU for a response nobody will use. An end-to-end timeout budget gives the request one bounded lifetime and lets each hop consume part of it.

Cloud Computing 01 Sep 2026 5 min read

Idempotent Event Consumers for At-Least-Once Delivery

Many queues and event brokers provide at-least-once delivery: a message that has been accepted can be delivered again when acknowledgements are lost, consumers crash, visibility timeouts expire, or the broker retries after uncertain outcomes. Duplicates are therefore not exceptional. A robust consumer should assume that the same logical event can arrive more than once. Why duplicates happen Consider this sequence: a consumer receives an event; it updates the database successfully; the process crashes before acknowledging the message; the broker makes the message visible again; another consumer receives it. The broker cannot know that the database update happened. Redelivery is the safer choice.

Cloud Computing 01 Sep 2026 3 min read

Design Circuit Breakers for Cloud Service Dependencies

When a downstream service is failing, continuing to send every request can waste threads, sockets, and latency budget while increasing load on the unhealthy dependency. A circuit breaker temporarily fails calls fast after failures cross a threshold. Understand the three states A typical breaker is closed while calls flow normally. It moves open after a configured failure policy is exceeded. After a recovery interval, it becomes half-open and allows a limited number of probes. Successful probes close the circuit; failures open it again.

Software Engineering 01 Sep 2026 4 min read

Contract Tests for Reliable Service Boundaries

Distributed systems fail in an awkward place: each service can pass its own tests while the interaction between two services is incompatible. A provider may rename a JSON field, tighten validation, change an enum, or stop returning a value that a consumer quietly depends on. Contract tests make those cross-service assumptions executable. What a contract is A contract describes an observable interaction between a consumer and a provider. For an HTTP API, it might specify: