Skip to content

Archive

Reliability

108 articles
Software Engineering 21 Sep 2026 6 min read

Circuit Breakers Limit Cascading Failure

Circuit Breakers Limit Cascading Failure A slow or failing dependency can consume more than its own capacity. Callers wait, retry, hold sockets, occupy worker slots, and retain memory while requests remain unresolved. As pressure spreads upstream, a local fault can become a service-wide saturation event. A circuit breaker places a stateful gate around calls to that dependency. It observes outcomes, opens when the configured failure policy is met, rejects calls for a period, then permits a small number of probes. Successful probes can return the breaker to normal traffic; failed probes send it back to the open state.

Software Engineering 21 Sep 2026 7 min read

Circuit Breakers Bound Calls to an Unhealthy Dependency

Circuit Breakers Bound Calls to an Unhealthy Dependency A remote dependency can fail in a way that is both persistent and expensive. Connections time out, worker slots remain occupied, request queues grow, and retries add more traffic to a service that is already unable to respond. A caller that keeps issuing the same class of request can turn one dependency failure into pressure across its own process. A circuit breaker places a stateful admission decision in front of those calls. While the dependency behaves within policy, calls pass through. After enough qualifying failures, the breaker opens and rejects new calls locally for a period. Recovery is tested with limited traffic rather than a full return to normal load.

Software Engineering 21 Sep 2026 7 min read

Bulkheads Isolate Concurrency Across Failure Domains

Bulkheads Isolate Concurrency Across Failure Domains A service can remain reachable while its useful capacity disappears. A slow dependency holds requests open, those requests occupy workers or connection slots, and unrelated traffic waits behind work that cannot finish promptly. The fault began in one path, but a shared resource pool lets it consume capacity needed by every path. Bulkhead isolation partitions that finite capacity. Calls associated with one failure domain receive a bounded share rather than competing without separation for the entire pool. When one partition fills, admission fails or waits within that partition while capacity assigned to other work remains available.

Software Engineering 21 Sep 2026 6 min read

Bounded Queues Turn Overload into Explicit Backpressure

Bounded Queues Turn Overload into Explicit Backpressure A queue can absorb a short mismatch between arrival rate and processing rate. That buffering is useful when bursts are temporary. It becomes dangerous when the queue has no meaningful bound: sustained overload no longer appears as an admission failure, but as a growing backlog, rising memory use, and requests that finish long after their latency budget has expired. A bounded queue changes the contract. Once capacity is exhausted, the producer must wait, reject, shed, or route work elsewhere. The overload is no longer hidden inside an expanding buffer.

Software Engineering 21 Sep 2026 6 min read

Backpressure Bounds Work When Consumers Fall Behind

Backpressure Bounds Work When Consumers Fall Behind A fast producer and a slower consumer can coexist for a short burst if a buffer absorbs the difference. The same arrangement becomes unstable when the rate mismatch persists. Pending work accumulates, memory rises, latency stretches, and items may expire before a consumer reaches them. Backpressure changes the contract between both sides. Instead of accepting work without regard to downstream state, the system exposes limited capacity to the producer. When that capacity is exhausted, production pauses, admission is rejected, or another explicit overload policy takes effect.

Software Engineering 21 Sep 2026 7 min read

Adaptive Concurrency Limits Track Available Service Capacity

Adaptive Concurrency Limits Track Available Service Capacity A fixed concurrency ceiling is easy to operate when service capacity is stable. Real systems rarely stay in one operating regime. Database contention, cache hit rate, request mix, downstream latency, CPU availability, and deployment changes can all move the amount of work a service can sustain at once. Adaptive concurrency control treats the in-flight limit as a control variable. The limiter admits work up to a current ceiling, observes service behavior, then adjusts that ceiling. The aim is not maximum concurrency. It is enough concurrency to use available capacity without allowing queues to grow far beyond the useful operating region.

Software Engineering 20 Sep 2026 6 min read

Visibility Timeouts Turn Message Delivery into a Renewable Lease

A queue consumer often needs time to perform work before it can safely acknowledge a message. Removing the message at receive time would make a consumer crash capable of losing work. Keeping it immediately available would let several consumers process the same item at once. A visibility timeout occupies the middle ground. Receiving a message makes it temporarily unavailable to competing consumers. The consumer gets a bounded interval to finish and acknowledge it. If that interval expires first, the queue can expose the message for another delivery.

Software Engineering 20 Sep 2026 6 min read

Token Buckets Separate Sustained Rate from Burst Capacity

A rate limit expressed only as “100 requests per second” leaves an important policy question open. Can a client send 100 requests at the first instant of each second, or must those requests be spread evenly? A token bucket makes that distinction explicit by separating sustained rate from burst capacity. The limiter maintains a balance of tokens up to a fixed capacity. Tokens arrive at a configured refill rate. An operation is admitted only when enough tokens are available, and admission deducts its cost from the balance. Idle time accumulates capacity for a later burst, but never beyond the bucket limit.

Software Engineering 20 Sep 2026 6 min read

Stale-While-Revalidate Keeps Cache Refresh off the Request Path

Stale-While-Revalidate Keeps Cache Refresh off the Request Path A cache entry does not become useless at the exact instant its freshness timer expires. For some data, a value that is a few seconds old is still preferable to making every caller wait for a backend refresh. Stale-while-revalidate uses that tolerance explicitly: the cache may serve an expired value for a bounded interval while a refresh runs separately. The policy changes refresh from a request-path requirement into background work for entries that remain acceptable while stale. It can reduce latency spikes around expiration, but only when the application can state how stale a response may become.

Software Engineering 20 Sep 2026 7 min read

Load Shedding Protects Useful Work When Capacity Is Exhausted

A service can be healthy at 2,000 requests per second and collapse at 2,400. The extra 400 requests do not merely wait their turn. They may occupy connection slots, queue entries, memory, worker threads, database sessions, and retry budgets while useful throughput falls. Load shedding places an explicit admission decision before a scarce resource is fully consumed. When the system cannot serve all incoming work within its operating envelope, it rejects selected requests early instead of allowing every request to compete until they all become slow.

Software Engineering 20 Sep 2026 6 min read

Lease Renewal Needs a Safety Margin Before Expiry

Lease Renewal Needs a Safety Margin Before Expiry A lease grants a holder temporary authority until a recorded expiry. Keeping that authority requires renewal before the deadline. Scheduling renewal at the deadline itself leaves no room for network delay, scheduler pauses, storage latency, or a transient retry. A safer design attempts renewal earlier. The interval between the planned renewal and expiry is a safety margin: time reserved for ordinary uncertainty before the lease is treated as lost.

Software Engineering 20 Sep 2026 7 min read

Idempotency Keys Make Retried Writes Safe to Repeat

Idempotency Keys Make Retried Writes Safe to Repeat A client can lose the response to a successful write. The connection may close after the server commits a payment, creates an order, or schedules a job but before the response reaches the caller. From the client’s perspective, failure and success can look identical. Retrying blindly is dangerous for operations with non-idempotent effects. Sending the same POST twice may create two resources or charge twice. Refusing to retry leaves the caller with an ambiguous outcome.

Software Engineering 20 Sep 2026 6 min read

Hedged Requests Cut Tail Latency at a Controlled Cost

Hedged Requests Cut Tail Latency at a Controlled Cost A service can have acceptable median latency and still produce a small set of very slow responses. Queueing, runtime pauses, storage contention, packet loss, or a temporarily busy replica can push individual requests far beyond the common case. Hedged requests address that tail by sending a second copy after the first request has been outstanding for a chosen delay. The caller accepts the first valid response and cancels the remaining attempt. The technique trades a bounded amount of extra work for a chance to escape an unusually slow execution path.

Software Engineering 20 Sep 2026 7 min read

Deadline Propagation Stops Work After Callers Give Up

Deadline Propagation Stops Work After Callers Give Up A timeout at the edge does not automatically stop work deeper in a system. A client may abandon a request after two seconds while an API server continues waiting on another service, which may still be running a database query. The response has lost its consumer, yet CPU time, connections, memory, queue positions, and downstream capacity can remain occupied. Deadline propagation carries the caller’s time budget across those boundaries. Each component receives an absolute deadline or an equivalent remaining budget, refuses work that cannot start in time, and cancels operations when the budget expires. The goal is not merely faster failure. It is to keep useless work from surviving longer than the request that justified it.

Software Engineering 20 Sep 2026 6 min read

Deadline Propagation Stops Expired Requests from Consuming Downstream Capacity

A timeout placed only at the outer edge of a request does not automatically limit the work started deeper in the call graph. The client may stop waiting after 800 milliseconds while an internal service continues a database query, a remote call, or a queued task for several more seconds. The response is already useless to that client, yet the system is still spending capacity on it. Deadline propagation carries the request’s time boundary with the work. Each component can compare that boundary with its current clock, reserve time for its own processing, and refuse or cancel work that no longer fits. The result is not merely faster failure. It is a tighter relationship between useful work and resource consumption.

Software Engineering 20 Sep 2026 7 min read

Bulkheads Isolate Resource Pools Before Failures Spread

A service can remain healthy at the process level while becoming useless because one workload has consumed every scarce execution resource. A slow dependency can occupy all outbound connections. A noisy tenant can fill every worker slot. A background job can take the same semaphore permits needed by interactive requests. Bulkhead isolation limits that coupling. Instead of letting unrelated work compete for one undifferentiated pool, the system partitions selected resources and gives each class of work a bounded share. Saturation then has a smaller blast radius.

Software Engineering 20 Sep 2026 5 min read

Bulkheads Isolate Concurrency Across Dependencies

Bulkheads Isolate Concurrency Across Dependencies A service can have plenty of CPU and still become unavailable because one dependency stops completing work. Requests waiting on a slow database, remote API, or storage service retain execution slots, connections, memory, and queue positions. If unrelated operations share the same finite pool, one saturated path can consume capacity needed by healthy paths. Bulkhead isolation divides that shared concurrency into explicit budgets. Calls to one dependency or workload class use a bounded pool that other classes cannot exhaust. The pattern does not repair a failing dependency. It limits the amount of local capacity that failure can occupy.

Software Engineering 20 Sep 2026 7 min read

Bounded Queues Turn Overload into Explicit Rejection

Bounded Queues Turn Overload into Explicit Rejection A queue can absorb short bursts when requests arrive faster than workers can finish them. That buffer is useful only while it remains a buffer. If producers can keep adding work without a fixed limit, sustained overload turns the queue into an expanding inventory of requests that may wait long after their results are useful. A bounded queue changes the failure mode. It accepts waiting work up to a deliberate capacity, then refuses additional admission until space becomes available. The service still experiences overload, but the overload appears as an explicit control decision rather than unbounded growth in memory and waiting time.

Software Engineering 20 Sep 2026 6 min read

Backpressure Keeps Producer Speed Tied to Consumer Capacity

A fast producer and a slower consumer can coexist safely only while the gap between their rates remains bounded. If incoming work arrives faster than it can be completed for long enough, buffering does not remove overload. It stores the difference. Backpressure makes that capacity mismatch part of the protocol between components. Instead of accepting work indefinitely, a saturated stage causes upstream code to slow down, wait for capacity, reduce demand, or reject work according to an explicit policy.

Software Engineering 20 Sep 2026 7 min read

Adaptive Concurrency Limits Follow Service Capacity

Adaptive Concurrency Limits Follow Service Capacity A service can become slower before it becomes unavailable. As in-flight work rises, CPU queues grow, connection pools fill, lock contention increases, and downstream calls accumulate. A fixed concurrency ceiling can protect the service, but one number rarely fits every operating condition. Capacity shifts with request mix, cache hit rate, dependency latency, deployment shape, and resource pressure. Adaptive concurrency control treats the admission limit as a value that can move. The controller observes recent service behavior, raises the limit while additional concurrency remains productive, and reduces it when latency indicates growing queues or saturation. The goal is not maximum concurrency. It is enough parallel work to use available capacity without allowing queues to dominate response time.

Tech 16 Sep 2026 6 min read

TCP TIME-WAIT Preserves Closed Connection State

A TCP endpoint can finish an application’s close operation while the protocol still retains state for that connection. After an active close completes its FIN exchange, the endpoint normally enters TIME-WAIT instead of discarding the connection record immediately. That retained state has two jobs. It leaves the endpoint able to acknowledge a retransmitted final FIN, and it separates a closed connection from a later incarnation that could use the same local and remote addresses and ports.

Software Engineering 16 Sep 2026 5 min read

TCP TIME-WAIT Delays Four-Tuple Reuse

A TCP endpoint that performs the active close can keep the closed connection in TIME-WAIT after the final ACK has been sent. The application-visible stream is finished, yet the transport retains state for a bounded interval before permitting unrestricted reuse of the same connection identity. That retention is not leftover application state. It protects the protocol boundary between one connection incarnation and a later connection that could otherwise use the same source address, source port, destination address, and destination port.

Software Engineering 15 Sep 2026 7 min read

Idempotency Keys Bind Retries to One Logical Operation

A client can transmit a state-changing request, lose the response, and retry without knowing whether the first attempt committed. At that boundary, transport failure has created ambiguity rather than proof of application failure. Repeating the mutation blindly can create a second logical effect. An idempotency key gives the server a stable operation identity across those delivery attempts. The first accepted request associates the key with an operation record. A later request carrying the same key can then reuse the recorded outcome instead of executing the mutation again.

Software Engineering 15 Sep 2026 6 min read

Half-Open TCP Connections Hide Peer Failure Until Traffic Resumes

A TCP socket can remain in the established state on one host after the peer has become unreachable or has lost all connection state. No contradiction exists in that state: TCP endpoints maintain local protocol state, and a silent network failure does not automatically deliver evidence that the peer is gone. This creates a boundary between connection state and peer liveness. An established socket records what the local TCP implementation currently knows about a byte-stream association. It is not a continuously refreshed assertion that the remote process, host, route, and intervening network are all operational.