Skip to content

Archive

Reliability

108 articles
Software Engineering 09 Sep 2026 11 min read

Using Hedged Requests to Reduce Tail Latency

Most calls to a dependency may finish quickly while a small fraction take much longer. A request that depends on one of those slow calls inherits the delay even when another healthy instance could have answered sooner. Increasing the timeout does not solve this problem. Retrying only after the timeout may also be too late: by then, the caller has already spent most of its latency budget. A hedged request is a deliberately delayed duplicate of an operation that is still in progress. The original request starts normally. If it has not completed after a chosen delay, the caller sends one additional equivalent request, usually to another eligible instance. The first acceptable result wins, and the remaining work is cancelled or ignored.

Software Engineering 09 Sep 2026 9 min read

Transactional Outbox for Reliable Message Publishing

Transactional Outbox for Reliable Message Publishing A common service operation has to do two things: change its own data and tell another part of the system what happened. For example, an order service may mark an order as paid and publish an OrderPaid message. The awkward part is that the database and message broker usually have separate commit mechanisms. If the service updates the database and then publishes, it can crash between those steps. If it publishes first, the database update can fail afterward. Either order can leave the two systems disagreeing about what happened.

Software Engineering 09 Sep 2026 11 min read

Propagating Deadlines Through Call Chains

A request can have a timeout at every network call and still take far longer than the caller intended. The problem appears when each layer starts a fresh timeout. A frontend gives service A 800 milliseconds. Service A spends 300 milliseconds doing local work, then gives service B another 800 milliseconds. Service B spends 250 milliseconds and gives service C yet another 800 milliseconds. Every individual timeout looks reasonable, but the chain no longer has an 800-millisecond limit.

Software Engineering 09 Sep 2026 9 min read

Designing Operations to Avoid Partial State

A function can report an error and still leave trouble behind. The first few steps may have changed state before a later step failed, so the caller receives a failure while the system now contains a mixture of old and new values. This is partial state: an operation did not complete, but some of its intended changes became visible. Partial state makes retrying, debugging, and reasoning about invariants harder because “the operation failed” no longer tells you what state remains.

Software Engineering 09 Sep 2026 9 min read

Containing Failures with Bulkheads

A service can fail even when most of its dependencies are healthy. One slow dependency may occupy every worker, connection, or concurrency slot until unrelated requests can no longer make progress. This is a resource-isolation problem. The dependency failure matters, but the larger outage happens because the system lets one workload consume capacity that other workloads also need. A bulkhead limits that sharing. It gives a workload a bounded resource budget so trouble in one area is less able to exhaust resources needed elsewhere. This article explains the mental model, shows where bulkheads help, and covers the trade-offs that make isolation useful rather than arbitrary.

Software Engineering 09 Sep 2026 10 min read

Coalescing Duplicate In-Flight Work

A service can receive many requests for the same expensive result at almost the same time. If every request starts identical work, a brief traffic burst can become a much larger burst against a database, remote API, filesystem, or CPU-heavy computation. Caching can help after a result exists. It does not necessarily help when the result is missing and many callers discover that miss together. Request coalescing solves this narrower problem. While an operation for a particular key is already running, later callers for the same key join that operation instead of starting another one. When it finishes, the waiting callers receive the same outcome.

Software Engineering 08 Sep 2026 9 min read

Testing Resilience with Fault Injection

A service can pass every normal-path test and still behave badly when a dependency times out, a write fails halfway through, or a connection disappears at an inconvenient moment. The problem is often not missing error handling. It is that the team has never observed whether the error handling produces the behavior they expect. Fault injection is the deliberate introduction of a controlled failure into a system or test. Instead of waiting for a real dependency to fail, you make a specific failure happen and observe the consequence.

Software Engineering 08 Sep 2026 10 min read

Metamorphic Testing When Exact Answers Are Hard

Some software is easy to test because the expected answer is obvious. If a function adds two numbers, a test can call it with 2 and 3 and assert that the result is 5. Other software has outputs that are expensive or awkward to predict. A route planner may examine thousands of possible paths. A search ranker may score hundreds of candidates. A numerical routine may produce a result that is difficult to calculate independently without reimplementing the same algorithm.

Software Engineering 08 Sep 2026 10 min read

Making Resource Ownership Explicit

A function opens a file and returns a parser. A factory creates a client backed by a connection pool. A component starts a worker and hands another component a handle. Everything works until shutdown, an exception, or a refactor exposes a basic question that the design never answered: who is responsible for releasing the resource? Resource leaks are often described as missing cleanup calls. That is only the visible failure. The deeper design problem is ambiguous ownership. If several parts of the program can use a resource but none clearly owns its lifetime, cleanup becomes guesswork.

Software Engineering 08 Sep 2026 11 min read

Designing Read-Your-Writes Consistency

A user changes their delivery address, sees a success message, and opens the order page. The old address appears. They refresh a few seconds later and the new value finally shows up. Nothing necessarily lost the write. The system may have accepted it correctly and then served the next read from a copy that had not caught up yet. From the user’s perspective, however, a successful update appeared to reverse itself.

Software Engineering 07 Sep 2026 10 min read

Using Error Budgets to Balance Reliability and Change

Teams often agree that reliability matters and still struggle to decide what to do when reliability competes with product work. Should a risky release proceed? Should engineers stop feature work after a bad incident? Is one failed request enough to justify a freeze? Without a shared rule, these decisions can become arguments between vague goals: “move faster” versus “make it more reliable.” An error budget turns a reliability target into a limited allowance for unsuccessful service. The budget does not make failures desirable. It makes the acceptable amount of unreliability explicit so a team can reason about risk and change using the same constraint.

Software Engineering 07 Sep 2026 11 min read

Understanding Two-Phase Commit

Suppose one business operation must update two independent transactional resources. Writing to the first and then the second creates an uncomfortable failure case: the first write may commit while the second fails. Reversing the order only moves the problem. Two-phase commit, usually shortened to 2PC, is a coordination protocol for making one commit-or-abort decision across multiple participants that can each prepare and commit a local transaction. Its purpose is atomicity across those participants: under the protocol’s assumptions, they do not intentionally finish with some participants committed and others aborted for the same transaction.

Software Engineering 07 Sep 2026 9 min read

Testing Failure Semantics, Not Just Errors

A test that expects an error can still miss the most damaging part of a failure. An order operation may return payment declined correctly while also marking the order as paid. A file import may report that parsing failed after writing half of its records. A retry may succeed but send the same notification twice. In each case, the visible error is correct. The failure semantics are not. Failure semantics describe what the system promises about state, side effects, and subsequent operations when something goes wrong. Testing them means asking more than “did this fail?” This article shows how to identify those promises and turn them into focused tests.

Software Engineering 07 Sep 2026 9 min read

Property-Based Testing for Behavioral Invariants

Example-based tests are excellent when you know the cases that matter. You choose an input, state the expected result, and protect that behavior from regression. The weakness is also clear: the test checks only the examples you thought to write. Some defects hide between those examples. A parser works for ordinary names but fails on an empty segment. A range-normalization function works for the three values in the test file but produces an invalid range for an unusual ordering. A serializer handles familiar records but loses information for one combination of optional fields.

Software Engineering 07 Sep 2026 9 min read

Evolving Data Shapes with Expand and Contract

A data change can look trivial in code and still be dangerous in a running system. Renaming a field, splitting one value into two, or changing a message shape may require several application versions, background jobs, and consumers to coexist while the change is in progress. The risky assumption is that the whole system changes at once. In practice, deployments take time, workers may finish old jobs after new code is live, and independently deployed consumers may upgrade later. If one release removes the old shape while something still depends on it, a locally correct change becomes a system failure.

Software Engineering 07 Sep 2026 11 min read

Circuit Breakers for Failing Dependencies

A dependency can fail in a way that is worse than an immediate error. It may accept connections but respond slowly, time out repeatedly, or reject nearly every request while callers continue sending more work. If your application keeps calling that dependency for every incoming request, the original failure can consume connection pools, worker capacity, and latency budgets in your own service. Retrying every failure can increase the pressure further.

Software Engineering 06 Sep 2026 9 min read

Designing APIs with Preconditions and Postconditions

An API can have clear parameter names and still leave its most important rules unstated. Can a withdrawal amount be zero? Must an account already be open? If a call succeeds, is the balance guaranteed to have changed, or has the request merely been accepted for later processing? When those questions are unclear, callers make assumptions. Different callers may make different assumptions, and failures appear far from the decision that caused them.

Software Engineering 06 Sep 2026 11 min read

Design Retry Budgets to Prevent Retry Storms

Retries are one of the simplest reliability tools in distributed systems. A transient network failure, a short leader election, or a momentary overload can make a request fail even though the dependency becomes healthy again a few milliseconds later. Retrying can hide that temporary failure from the user. The same mechanism can also make an outage worse. If a struggling dependency starts failing requests and every caller retries immediately, the dependency receives extra work precisely when it has the least capacity to handle it. A small failure rate can turn into a retry storm: retries create more load, more load creates more failures, and those failures create still more retries.

Software Engineering 05 Sep 2026 9 min read

Designing Idempotent Operations for Safe Retries

A caller sends a request to create a payment. The server processes it, but the response is lost because the connection closes. The caller now has a difficult choice: retry and risk charging twice, or stop and risk leaving the payment incomplete. This is not mainly a networking problem. It is an operation-design problem. When a caller cannot tell whether an attempt succeeded, retrying is only safe when the system has a way to recognize that the new attempt represents the same intent.

Linux 04 Sep 2026 8 min read

Write Files Atomically on Linux with rename and fsync

Updating a small file looks simple: open it, truncate it, write the new contents, and close it. That works when nothing interrupts the write. The failure mode appears when a process crashes, the machine loses power, or another process reads the file while it is being rewritten. A reader can observe an empty or partially written file, and a crash can leave the pathname referring to incomplete data. For configuration files, state snapshots, generated metadata, and similar single-file updates, a better pattern is to write a complete replacement beside the original file and then rename it into place.

Software Engineering 04 Sep 2026 10 min read

Designing Failure Containment Boundaries

A component fails. Soon unrelated requests become slow, worker queues stop moving, and healthy features begin returning errors. The original defect may be small, but the system has allowed its effects to spread. Failure containment is the design practice of limiting how far a fault can propagate. The goal is not to prevent every failure. That is unrealistic. The goal is to make a local failure stay local enough that the rest of the system can continue useful work or fail in a controlled way.

Cloud Computing 04 Sep 2026 11 min read

Design Dead-Letter Queues for Poison Messages

Retries are useful when a failure is temporary. A database may be unavailable for a few seconds, a downstream service may return an overload response, or a network connection may disappear and recover. Retries become harmful when the message itself cannot succeed. A malformed payload, an unsupported schema version, a reference to permanently missing data, or a deterministic application bug can make the same message fail on every delivery. If the broker keeps returning that message to consumers indefinitely, the system spends capacity repeating work that has no chance of succeeding.

Software Engineering 04 Sep 2026 9 min read

Assertions That Expose Broken Assumptions

A program can produce the wrong result long after the code that caused the problem has run. A function corrupts an internal value, several operations accept it, and an unrelated component eventually fails. By then, the stack trace points at the consequence rather than the cause. Assertions help shorten that distance. An assertion states an internal condition that the programmer believes must be true at a particular point in the program. If the condition is false, the program reports a broken assumption immediately instead of continuing as though the state were valid.

Software Engineering 03 Sep 2026 8 min read

Designing Errors for Actionable Failures

Errors are part of a program’s interface. They tell callers that an operation could not produce its promised result, but useful error handling goes further: it preserves enough meaning for the caller to decide what to do next. Weak error handling tends to fail in two opposite ways. Some code hides failures by returning defaults, logging and continuing, or catching exceptions too broadly. Other code exposes every low-level detail directly, forcing callers to understand implementation choices that should have remained private.