Skip to content

Archive

Reliability

108 articles
Software Engineering 22 Sep 2026 6 min read

Transactional Outbox Closes the Database-to-Broker Commit Gap

Transactional Outbox Closes the Database-to-Broker Commit Gap A service often needs one operation to change database state and emit a message. An order may become confirmed while an OrderConfirmed event is sent to a broker. Those actions touch separate systems, so two ordinary writes cannot form one atomic commit unless both systems participate in a distributed transaction. The dangerous part is the interval between the writes. Commit the database first and the process can fail before publishing. Publish first and the database commit can fail afterward. Reversing the order moves the failure window; it does not remove it.

Software Engineering 22 Sep 2026 7 min read

Request Coalescing Stops Cache Misses from Multiplying Backend Work

Request Coalescing Stops Cache Misses from Multiplying Backend Work A cache miss is usually cheap when one caller causes one backend lookup. The same miss can become expensive when many callers arrive for the same key at nearly the same time. Each caller observes the key as absent, each starts identical work, and the backend receives a burst precisely when the cache is providing no protection for that key. Request coalescing changes that concurrency pattern. The first caller becomes the leader for a key. Later callers join the same in-flight operation and wait for its result rather than starting equivalent work. Once the fill completes, the result can populate the cache and be returned to the waiting callers.

Software Engineering 22 Sep 2026 8 min read

Optimistic Concurrency Rejects Stale Writes Before They Replace Newer State

Optimistic Concurrency Rejects Stale Writes Before They Replace Newer State A read-modify-write flow looks harmless when only one actor touches a record. A client reads state, changes part of it, then writes the result back. With concurrent actors, the interval between the read and the write becomes a race. Another writer can commit a newer value during that interval, and an unconditional update can erase it. Optimistic concurrency control puts a condition on the final write. The client carries a version derived from the state it read, and the storage layer accepts the mutation only if that version is still current. A mismatch becomes a conflict rather than a silent overwrite.

Software Engineering 22 Sep 2026 5 min read

Load Shedding Protects Services When Capacity Runs Out

Load Shedding Protects Services When Capacity Runs Out A service can receive more work than it can complete. The first visible symptom is often not an immediate error but a queue that grows while workers remain fully occupied. Requests spend longer waiting, deadlines expire, clients retry, and the extra retry traffic can deepen the overload. Load shedding places an explicit rejection point before that spiral consumes every available resource. The service admits work that fits its operating capacity and fails excess work quickly enough to preserve useful throughput for requests that can still complete.

Software Engineering 22 Sep 2026 6 min read

Idempotency Keys Turn Retries into One Logical Operation

Idempotency Keys Turn Retries into One Logical Operation A client can lose the response to a successful request. The server may commit a charge, create an order, or enqueue a job, then the connection can fail before the response reaches the caller. From the client’s view, success and failure are now ambiguous. Retrying is necessary for availability, but an ordinary retry can repeat the side effect. An idempotency key gives the client a way to say that several HTTP attempts represent one logical operation.

Software Engineering 22 Sep 2026 8 min read

Hedged Requests Trade Duplicate Work for Lower Tail Latency

Hedged Requests Trade Duplicate Work for Lower Tail Latency Most requests may finish quickly while a small fraction take much longer. A busy worker, a transient network queue, a cold cache entry, garbage collection, storage contention, or another local disturbance can stretch one attempt far beyond the median. At scale, those slow outliers become visible in p95, p99, and higher-percentile latency even when average service time looks healthy. A hedged request sends an additional attempt after the original has been outstanding for a chosen delay. Both attempts represent the same logical operation. The caller accepts the first valid result and cancels or ignores the remaining attempt.

Software Engineering 22 Sep 2026 6 min read

Fencing Tokens Stop Stale Lease Holders

Fencing Tokens Stop Stale Lease Holders A distributed lease gives one worker temporary permission to act as an owner. The lease eventually expires so another worker can take over after a crash or network failure. That solves availability, but expiry alone does not guarantee that the old worker has stopped. A process can pause long enough for its lease to expire, then resume with stale local state. A long garbage-collection pause, scheduler stall, suspended virtual machine, or delayed network path can create this condition. If the old worker writes after a replacement has taken ownership, two workers can affect the same resource even though the lease service never considered both leases valid at the same instant.

Software Engineering 22 Sep 2026 6 min read

Dead-Letter Queues Isolate Poison Messages Without Blocking Progress

Dead-Letter Queues Isolate Poison Messages Without Blocking Progress A message consumer usually treats failure as temporary at first. A database may be unavailable, a remote service may time out, or a worker may restart between receiving and acknowledging a message. Retrying is appropriate when another attempt has a reasonable chance of succeeding. Some messages fail for a different reason. Their payload is malformed, a referenced entity can never satisfy a required condition, or the consumer has a deterministic defect triggered by that input. Repeated delivery then consumes capacity without moving the message toward completion. A dead-letter queue gives that failure a separate destination after the normal retry policy is exhausted.

Software Engineering 22 Sep 2026 7 min read

Circuit Breakers Stop Repeated Calls to Failing Dependencies

Circuit Breakers Stop Repeated Calls to Failing Dependencies A dependency that is already failing can consume more caller capacity than a healthy one. Requests wait for timeouts, retries add traffic, connection pools remain occupied, and worker slots stay tied to work that has little chance of completing. A circuit breaker places a stateful gate in front of that dependency so the caller can stop issuing calls after failure evidence reaches a configured limit.

Software Engineering 22 Sep 2026 6 min read

Bulkheads Keep One Saturated Dependency from Consuming Every Worker

Bulkheads Keep One Saturated Dependency from Consuming Every Worker A service can have healthy CPU, available memory, and responsive internal code while still becoming unavailable. One downstream dependency is enough to consume the service’s entire concurrency budget if calls to it become slow and every request is allowed to wait. The failure is not limited to the slow dependency. Shared worker pools, connection pools, semaphores, queues, and request slots turn local saturation into a service-wide outage. Bulkhead isolation limits that blast radius by reserving separate capacity for distinct workloads or dependencies.

Software Engineering 22 Sep 2026 5 min read

Bulkheads Isolate Concurrency Before One Dependency Consumes It All

Bulkheads Isolate Concurrency Before One Dependency Consumes It All A service can have enough CPU and memory yet stop making useful progress because its concurrency is exhausted. Threads, database connections, outbound sockets, worker slots, and in-flight request permits are finite. If one dependency becomes slow, calls to that dependency can occupy the entire shared pool. Bulkhead isolation divides that capacity before saturation occurs. Workloads that can fail independently receive separate concurrency budgets, so pressure in one path does not automatically consume every slot needed by another.

Software Engineering 22 Sep 2026 7 min read

Bounded Queues Turn Overload into an Explicit Admission Decision

Bounded Queues Turn Overload into an Explicit Admission Decision A queue absorbs short differences between arrival rate and service rate. That buffer is useful when a burst ends before workers fall far behind. The same mechanism becomes dangerous when arrivals remain faster than completions: every accepted item adds waiting time and consumes some combination of memory, descriptors, references, or durable storage. A bounded queue places a finite limit on that waiting population. Once the limit is reached, the system must make an admission decision instead of silently extending the backlog. Depending on the interface, that decision may block a producer, reject new work, shed selected work, or redirect it to another capacity domain.

Software Engineering 22 Sep 2026 6 min read

Backpressure Keeps Fast Producers from Overrunning Slow Consumers

Backpressure Keeps Fast Producers from Overrunning Slow Consumers A pipeline is stable only while work leaves each stage at roughly the rate it arrives over a useful time window. When a producer can submit work faster than a consumer can finish it, the difference has to accumulate somewhere. An unbounded queue makes that accumulation easy to miss. Requests continue to be accepted, the producer appears healthy, and the consumer keeps working. Meanwhile queued work consumes memory and ages before execution. Backpressure turns downstream saturation into an upstream signal before the backlog becomes the failure.

Software Engineering 21 Sep 2026 5 min read

Transactional Outbox Closes the Database-Broker Commit Gap

Transactional Outbox Closes the Database-Broker Commit Gap A service often needs one request to change database state and publish a message. Those actions may look adjacent in application code, but they cross two independent commit boundaries. If the database and broker do not share a transaction protocol, no ordering of two ordinary writes can make them atomic. Consider an order service that stores an accepted order and emits OrderCreated. Publishing after the database commit leaves a crash window before the broker call. Publishing first creates the opposite window: consumers can receive an event for state that later fails to commit.

Software Engineering 21 Sep 2026 6 min read

Token Buckets Preserve Burst Capacity Without Removing Rate Bounds

Token Buckets Preserve Burst Capacity Without Removing Rate Bounds A fixed requests-per-second ceiling treats a brief spike and a sustained flood as the same event. That can be too rigid for services whose callers naturally arrive in clusters. A token bucket separates two constraints: the long-run admission rate and the amount of burst traffic the service is willing to absorb. The model has two parameters. The bucket capacity B is the maximum number of tokens that can accumulate. The refill rate r adds tokens per unit of time, up to B. A request consumes tokens according to its configured cost. If enough tokens are present, the request proceeds; otherwise it is rejected, delayed, or handled by another explicit policy.

Software Engineering 21 Sep 2026 6 min read

Request Coalescing Collapses Cache-Miss Bursts

Request Coalescing Collapses Cache-Miss Bursts A cache miss is usually cheap when one caller triggers one backend read. The same miss can become expensive when hundreds of callers arrive for the same key before the first fill completes. Each caller sees an empty cache and starts equivalent work, multiplying load precisely when the cached value is unavailable. Request coalescing places a small concurrency boundary around that fill. The first caller starts the backend operation. Later callers for the same key join the in-flight operation instead of starting another one. When it completes, the result can populate the cache and be delivered to the waiting callers.

Software Engineering 21 Sep 2026 7 min read

Load Shedding Protects Useful Work Under Saturation

Load Shedding Protects Useful Work Under Saturation A service has a finite amount of work it can complete per unit of time. When offered load rises past that capacity, accepting every request does not create more capacity. It creates more waiting, consumes memory and connection slots, extends deadlines, and can leave expensive work running after callers have already given up. Load shedding makes admission explicit. Work that the service cannot process within its operating envelope is rejected early so admitted work retains a realistic chance of completing.

Software Engineering 21 Sep 2026 7 min read

Idempotency Keys Make Retried Mutations Safe

Idempotency Keys Make Retried Mutations Safe A client can lose the result of a successful mutation without losing the mutation itself. The server may commit a payment, reservation, or job submission and then drop the connection before the response reaches the caller. From the client’s perspective, a timeout leaves two plausible states: the operation failed before commit, or it committed and only the response was lost. Blindly retrying a non-idempotent mutation can apply the effect twice. Refusing every retry leaves the caller unable to recover from an ambiguous outcome. An idempotency key gives both sides a stable identity for one logical operation, so a repeated attempt can reuse the result of the first accepted attempt instead of creating another effect.

Software Engineering 21 Sep 2026 6 min read

Idempotency Keys Bound Duplicate Mutations Across Retries

Idempotency Keys Bound Duplicate Mutations Across Retries A client can lose the response to a successful mutation. The server may commit a payment, reservation, or job submission and then lose the connection before the response reaches the caller. From the client side, timeout does not reveal whether the mutation failed before commit or succeeded before the response disappeared. A retry is necessary for availability, but a blind retry can repeat the side effect. An idempotency key gives both attempts a stable identity so the server can treat them as one logical operation.

Software Engineering 21 Sep 2026 7 min read

Hedged Requests Cut Tail Latency with Controlled Duplication

Hedged Requests Cut Tail Latency with Controlled Duplication A service can have a healthy median latency while a small fraction of requests take far longer. Queueing, garbage collection, storage stalls, packet loss, noisy neighbors, or uneven replica load can all stretch the slow end of the distribution. For a request that fans out to several dependencies, one slow branch can dominate the entire response. Hedged requests reduce that exposure by starting a second copy after a short delay. The copies target independent execution paths when possible, and the first valid response wins.

Software Engineering 21 Sep 2026 6 min read

Hedged Requests Cut Tail Latency at a Capacity Cost

Hedged Requests Cut Tail Latency at a Capacity Cost A service can have acceptable median latency while a small fraction of requests take much longer. Queueing, a cold cache, runtime pauses, transient packet loss, or a slow storage operation can leave one attempt far behind the normal path. At sufficient fan-out, those rare delays become common at the aggregate request boundary. A hedged request starts a second equivalent attempt after the first has remained incomplete for a selected delay. The caller accepts the first useful result and cancels or discards the other attempt.

Software Engineering 21 Sep 2026 6 min read

Fencing Tokens Stop Stale Lease Holders from Writing

Fencing Tokens Stop Stale Lease Holders from Writing A distributed lease can decide which client currently owns a resource, but lease expiry does not instantly stop the previous holder. A process can pause, lose network access, or stall long enough for its lease to expire. Another client then acquires the lease. If the old process resumes and still has access to the protected storage or service, both clients can issue writes.

Software Engineering 21 Sep 2026 5 min read

Deadline Propagation Preserves Timeout Budgets Across RPC Hops

Deadline Propagation Preserves Timeout Budgets Across RPC Hops A timeout that restarts at every service boundary can turn a short caller budget into a much longer chain of work. A client may allow 800 milliseconds, service A may spend 500 milliseconds locally, then call service B with a fresh 800-millisecond timeout. B can continue working long after the client has stopped waiting. Deadline propagation keeps one end time attached to the request. Each hop derives its remaining budget from that deadline and refuses to start work that cannot fit within it. The result is not faster execution by itself. It is bounded execution that respects the time contract established upstream.

Software Engineering 21 Sep 2026 6 min read

Deadline Propagation Preserves Request Time Budgets

Deadline Propagation Preserves Request Time Budgets A request can cross several services before producing a response. Each hop may have its own queue, network call, retry policy, and local timeout. If those limits are chosen independently, the total path can run far longer than the caller is prepared to wait. Deadline propagation gives the path one temporal boundary. The initiating caller supplies an absolute deadline, or a time budget that is converted into one. Each downstream component uses the remaining interval rather than starting a fresh timeout from zero.