Skip to content

Archive

Resilience

15 articles
Software Engineering 22 Sep 2026 7 min read

Circuit Breakers Stop Repeated Calls to Failing Dependencies

Circuit Breakers Stop Repeated Calls to Failing Dependencies A dependency that is already failing can consume more caller capacity than a healthy one. Requests wait for timeouts, retries add traffic, connection pools remain occupied, and worker slots stay tied to work that has little chance of completing. A circuit breaker places a stateful gate in front of that dependency so the caller can stop issuing calls after failure evidence reaches a configured limit.

Software Engineering 21 Sep 2026 6 min read

Circuit Breakers Limit Cascading Failure

Circuit Breakers Limit Cascading Failure A slow or failing dependency can consume more than its own capacity. Callers wait, retry, hold sockets, occupy worker slots, and retain memory while requests remain unresolved. As pressure spreads upstream, a local fault can become a service-wide saturation event. A circuit breaker places a stateful gate around calls to that dependency. It observes outcomes, opens when the configured failure policy is met, rejects calls for a period, then permits a small number of probes. Successful probes can return the breaker to normal traffic; failed probes send it back to the open state.

Software Engineering 21 Sep 2026 7 min read

Circuit Breakers Bound Calls to an Unhealthy Dependency

Circuit Breakers Bound Calls to an Unhealthy Dependency A remote dependency can fail in a way that is both persistent and expensive. Connections time out, worker slots remain occupied, request queues grow, and retries add more traffic to a service that is already unable to respond. A caller that keeps issuing the same class of request can turn one dependency failure into pressure across its own process. A circuit breaker places a stateful admission decision in front of those calls. While the dependency behaves within policy, calls pass through. After enough qualifying failures, the breaker opens and rejects new calls locally for a period. Recovery is tested with limited traffic rather than a full return to normal load.

Software Engineering 20 Sep 2026 5 min read

Hedged Requests Trade Extra Work for Lower Tail Latency

A service can have a healthy median latency and still produce occasional requests that take far longer than the rest. Queueing, a slow replica, connection setup, garbage collection, storage stalls, or transient network delay can leave one attempt behind while equivalent capacity elsewhere remains available. A hedged request limits exposure to that single slow path. The client starts one attempt normally. If it is still pending after a configured delay, the client may start a second equivalent attempt. The first acceptable result is used, and the remaining attempt is cancelled when cancellation is supported.

Software Engineering 20 Sep 2026 4 min read

Circuit Breakers Limit Repeated Calls to Failing Dependencies

Circuit Breakers Limit Repeated Calls to Failing Dependencies A remote dependency can fail in a way that is both slow and expensive. Requests wait for timeouts, workers remain occupied, retries add more traffic, and a local service can lose capacity even when its own code is healthy. A circuit breaker places a stateful decision in front of that call path. While the dependency behaves acceptably, calls pass through. After the configured failure condition is reached, the breaker opens and rejects new calls locally for a bounded period. Later, it admits a small number of probes before deciding whether normal traffic can resume.

Software Engineering 19 Sep 2026 7 min read

Circuit Breakers Bound Failure Traffic Across Service Calls

A circuit breaker changes the admission decision for an outbound call before the dependency receives it. In the closed state, calls proceed and their outcomes feed a failure policy. Once that policy trips, the breaker enters the open state and rejects subsequent calls locally. After a configured recovery interval, a limited set of probe calls can test whether the dependency is usable again. That mechanism is distinct from retries. A retry issues another attempt after a failed attempt. A breaker can prevent an attempt from being issued at all. Combining the two without a precise ordering can amplify traffic during an outage or keep a breaker open based on signals that do not represent dependency health.

Software Engineering 19 Sep 2026 6 min read

Circuit Breakers Bound Failure Amplification Across Service Calls

A service call can fail quickly and still create a larger system problem. When every upstream request continues to invoke a downstream dependency that is already failing, each attempt consumes connection capacity, worker time, retry budget, and queue space. The dependency receives traffic it cannot currently serve, while callers spend resources waiting for outcomes that are already strongly correlated with recent failures. A circuit breaker puts a stateful decision boundary in front of that call. Instead of treating every request as an independent opportunity to try the dependency, it records recent failure state and can reject calls locally for a bounded interval. Recovery is then tested through controlled probes rather than a full return of traffic.

Software Engineering 13 Sep 2026 8 min read

Circuit Breakers Bound Failed Call Admission

A remote call can fail in a few milliseconds or consume its entire timeout budget before returning an error. If callers keep issuing equivalent requests while the dependency remains unable to serve them, each attempt spends resources on an outcome that recent evidence already suggests is unavailable. Retries can increase that pressure because one logical operation may create several physical calls. A circuit breaker changes call admission rather than the remote protocol. It records recent outcomes, moves between explicit states, and can reject new calls locally for a bounded period. After that period, it permits limited probes to test whether normal traffic can resume.

Software Engineering 12 Sep 2026 8 min read

Retry Amplification and the Role of Jitter

Retry Amplification and the Role of Jitter A failed request can create more traffic than a successful one. If a caller immediately repeats an operation after a transient error, the original unit of demand becomes two attempts. Add another retrying layer above that caller, and a single logical request can fan out into several physical attempts before any component has recovered. Retries are often described as a way to tolerate temporary faults. That description is incomplete because retry behavior also changes load. The mechanism sits inside a feedback loop: failure triggers another attempt, another attempt consumes capacity, and consumed capacity can affect the conditions that produced the failure.

Software Engineering 12 Sep 2026 9 min read

Circuit Breakers as Admission Control for Failing Dependencies

Circuit Breakers as Admission Control for Failing Dependencies A remote call that has little chance of succeeding still consumes something: a connection slot, a worker, a deadline budget, memory for request state, or capacity in the dependency itself. When repeated failures indicate that a downstream service is currently unable to serve useful work, continuing to admit every call can preserve the very pressure that callers need to escape. A circuit breaker changes that admission decision. Instead of treating each call as independent, it retains a small amount of state about recent outcomes. That state can temporarily reject new calls before network I/O begins, then permit controlled probes after a recovery interval.

Software Engineering 12 Sep 2026 11 min read

Bulkhead Isolation: Contain Failures with Separate Capacity Pools

Bulkhead Isolation: Contain Failures with Separate Capacity Pools A service can have enough total capacity and still become unavailable because one dependency consumes all of it. Imagine an API that calls a payment service and a recommendation service. Both outbound calls use the same worker pool. Recommendations become slow during a traffic spike. Their requests occupy every worker while waiting for responses. Payment requests now have no worker available, even though the payment service itself is healthy.

Software Engineering 07 Sep 2026 11 min read

Circuit Breakers for Failing Dependencies

A dependency can fail in a way that is worse than an immediate error. It may accept connections but respond slowly, time out repeatedly, or reject nearly every request while callers continue sending more work. If your application keeps calling that dependency for every incoming request, the original failure can consume connection pools, worker capacity, and latency budgets in your own service. Retrying every failure can increase the pressure further.

Software Engineering 06 Sep 2026 11 min read

Design Retry Budgets to Prevent Retry Storms

Retries are one of the simplest reliability tools in distributed systems. A transient network failure, a short leader election, or a momentary overload can make a request fail even though the dependency becomes healthy again a few milliseconds later. Retrying can hide that temporary failure from the user. The same mechanism can also make an outage worse. If a struggling dependency starts failing requests and every caller retries immediately, the dependency receives extra work precisely when it has the least capacity to handle it. A small failure rate can turn into a retry storm: retries create more load, more load creates more failures, and those failures create still more retries.

Cybersecurity 03 Sep 2026 5 min read

Design Backups for Security and Recovery

Backups are a security control, not only an operations convenience. They can limit the impact of ransomware, destructive mistakes, compromised administrator accounts, corrupted data, and failed deployments. A backup strategy is useful only when attackers cannot easily destroy it and the organization can reliably restore from it. Copying production data to another location is therefore only the beginning. Define what must be recoverable Start by identifying the systems and data that matter to recovery. This may include databases, uploaded files, configuration, encryption key material, infrastructure definitions, and other state that cannot simply be rebuilt from source control.

Cloud Computing 01 Sep 2026 3 min read

Design Circuit Breakers for Cloud Service Dependencies

When a downstream service is failing, continuing to send every request can waste threads, sockets, and latency budget while increasing load on the unhealthy dependency. A circuit breaker temporarily fails calls fast after failures cross a threshold. Understand the three states A typical breaker is closed while calls flow normally. It moves open after a configured failure policy is exceeded. After a recovery interval, it becomes half-open and allows a limited number of probes. Successful probes close the circuit; failures open it again.