Skip to content

Archive

Fault Tolerance

4 articles
Software Engineering 22 Sep 2026 7 min read

Fencing Tokens Block Stale Lock Holders

Fencing Tokens Block Stale Lock Holders A distributed lock is often used to keep two workers from changing the same resource at once. The difficult case begins when lock ownership depends on a lease. A client can acquire the lease, pause long enough for it to expire, then resume after another client has acquired a new lease. From the old client’s point of view, execution simply continued. From the coordination service’s point of view, ownership already moved. If the protected storage system accepts both clients’ writes, the old holder can overwrite work performed by the current holder.

Software Engineering 22 Sep 2026 5 min read

Bulkheads Isolate Concurrency Before One Dependency Consumes It All

Bulkheads Isolate Concurrency Before One Dependency Consumes It All A service can have enough CPU and memory yet stop making useful progress because its concurrency is exhausted. Threads, database connections, outbound sockets, worker slots, and in-flight request permits are finite. If one dependency becomes slow, calls to that dependency can occupy the entire shared pool. Bulkhead isolation divides that capacity before saturation occurs. Workloads that can fail independently receive separate concurrency budgets, so pressure in one path does not automatically consume every slot needed by another.

Software Engineering 20 Sep 2026 4 min read

Circuit Breakers Limit Repeated Calls to Failing Dependencies

Circuit Breakers Limit Repeated Calls to Failing Dependencies A remote dependency can fail in a way that is both slow and expensive. Requests wait for timeouts, workers remain occupied, retries add more traffic, and a local service can lose capacity even when its own code is healthy. A circuit breaker places a stateful decision in front of that call path. While the dependency behaves acceptably, calls pass through. After the configured failure condition is reached, the breaker opens and rejects new calls locally for a bounded period. Later, it admits a small number of probes before deciding whether normal traffic can resume.

Software Engineering 07 Sep 2026 11 min read

Circuit Breakers for Failing Dependencies

A dependency can fail in a way that is worse than an immediate error. It may accept connections but respond slowly, time out repeatedly, or reject nearly every request while callers continue sending more work. If your application keeps calling that dependency for every incoming request, the original failure can consume connection pools, worker capacity, and latency budgets in your own service. Retrying every failure can increase the pressure further.