Skip to content

Archive

Reliability

108 articles
Software Engineering 14 Sep 2026 7 min read

Poison Messages Turn Retries Into Queue Retention

A queue consumer receives a message, rejects it, and receives the same message again. That cycle is useful when the rejection came from a transient condition. It is structurally different when the payload can never be processed by the current consumer. The broker can keep honoring redelivery semantics while the application makes no forward progress on that message. Such a message is commonly called a poison message. The important property is not that it contains malformed bytes. A syntactically valid message can be permanently unprocessable because its schema is unsupported, a required invariant is violated, referenced data can never exist, or application logic deterministically rejects its state.

Software Engineering 14 Sep 2026 9 min read

Deadlines Shrink Across Service Boundaries

A service receives a request with 480 milliseconds remaining before its deadline. It spends 90 milliseconds reading state, then calls another service with a fixed 500-millisecond timeout. The downstream call can now outlive the request that caused it. Nothing about either timeout is internally inconsistent; the inconsistency appears at the boundary between them. Timeouts are often configured as local limits: a database query gets one value, an HTTP client another, a queue operation a third. A deadline represents a different constraint. It gives an operation an end point, so every later stage can compare its own work against the same finite lifetime.

Software Engineering 13 Sep 2026 7 min read

Monotonic Time Belongs in Elapsed-Time Measurement

A timeout can be represented by an ordinary timestamp comparison: record a start value, read the clock later, subtract, and compare the result with a limit. The arithmetic looks complete. The clock model is not. Civil time exists to place events on a shared calendar. It can be corrected to stay aligned with an external time reference. Elapsed-time measurement has a different requirement: later observations within one running system need an ordering suitable for measuring an interval. A clock correction that improves civil-time accuracy can therefore be harmful when the same reading is treated as a stopwatch.

Software Engineering 13 Sep 2026 9 min read

Cancellation Is a Protocol, Not a Thread Kill

A caller can stop waiting for a result while the operation producing that result continues to run. The distinction is easy to miss because many APIs expose cancellation through a single method, token, context, or signal. That surface can look like a command to terminate work. In most cooperative designs, it is closer to a state transition: further work is no longer wanted. The gap matters once an operation owns resources, crosses process boundaries, or has already produced side effects. A cancelled HTTP request does not retroactively erase a committed database transaction. A task that notices a cancellation token between two writes cannot make the first write disappear. A parent that abandons a child computation also needs a rule for who observes the child’s eventual completion and who releases anything the child owns.

Software Engineering 12 Sep 2026 12 min read

Saga Transactions: Coordinate Multi-Service Changes with Compensation

Saga Transactions: Coordinate Multi-Service Changes with Compensation A business operation can cross several services even when no single database transaction spans them all. An order flow might reserve inventory, authorize payment, create a shipment, and confirm the order. Each service owns its data and commits independently. That independence creates a difficult failure case. Inventory can be reserved successfully, then payment authorization can fail. A database rollback in the order service cannot undo a committed reservation in the inventory service.

Software Engineering 12 Sep 2026 11 min read

Load Shedding: Reject Work Before Overload Spreads

Load Shedding: Reject Work Before Overload Spreads A service has finite capacity. When offered work exceeds that capacity for long enough, accepting every request can make the service less useful rather than more useful. Queues grow, deadlines expire, memory pressure rises, dependencies receive more traffic, and successful throughput can fall. Load shedding is the deliberate rejection of work that the system cannot serve within an acceptable budget. The goal is not to maximize the number of requests admitted. The goal is to preserve useful service during overload.

Software Engineering 12 Sep 2026 11 min read

Idempotent Consumers: Handle Duplicate Messages Safely

Idempotent Consumers: Handle Duplicate Messages Safely A message broker can deliver the same message more than once. A worker may finish its database update and crash before acknowledging the message. The broker sees no acknowledgement, so it sends the message again. From the broker’s perspective, redelivery is the safe choice. From the application’s perspective, the second delivery can repeat a business effect. That gap matters whenever an effect must happen once per logical message. Charging an account twice, granting stock twice, incrementing a counter twice, or sending the same fulfillment request twice can turn a routine retry into corrupted state.

Software Engineering 12 Sep 2026 9 min read

Fencing Tokens: Block Stale Lease Holders

Fencing Tokens: Block Stale Lease Holders A distributed lease can grant one process temporary permission to act, but expiration alone cannot stop that process from acting after its lease has ended. A long pause, network delay, overloaded runtime, or suspended virtual machine can leave an old holder unaware that another process has already acquired the lease. This creates a subtle safety gap. Two processes can both believe they are entitled to modify the same resource, even when the lease service itself grants ownership correctly.

Software Engineering 12 Sep 2026 10 min read

Cache Stampede Control with Early Refresh

Cache Stampede Control with Early Refresh A cache can remove enormous amounts of repeated work, yet a popular entry creates a sharp risk at expiry. If ten thousand requests depend on the same key and that key expires, many requests can discover the miss at nearly the same moment. Each request may then start the same database query, computation, or remote call. This event is commonly called a cache stampede. The cache works well during the entry lifetime, then abruptly stops protecting the dependency exactly when demand is high.

Software Engineering 12 Sep 2026 11 min read

Bulkhead Isolation: Contain Failures with Separate Capacity Pools

Bulkhead Isolation: Contain Failures with Separate Capacity Pools A service can have enough total capacity and still become unavailable because one dependency consumes all of it. Imagine an API that calls a payment service and a recommendation service. Both outbound calls use the same worker pool. Recommendations become slow during a traffic spike. Their requests occupy every worker while waiting for responses. Payment requests now have no worker available, even though the payment service itself is healthy.

Software Engineering 12 Sep 2026 10 min read

Backpressure: Match Producer Speed to Consumer Capacity

Backpressure: Match Producer Speed to Consumer Capacity A pipeline is stable only when work enters at a rate its downstream stages can sustain. That sounds obvious, yet many systems let producers run at full speed until a queue fills, memory grows, latency explodes, or a downstream service starts rejecting requests. The visible failure appears late. The actual mismatch began earlier: one stage could create work faster than the next stage could finish it.

Software Engineering 11 Sep 2026 9 min read

Mutation Testing: Check Whether Tests Detect Broken Behavior

Mutation Testing: Check Whether Tests Detect Broken Behavior A test suite can execute every line of a function and still miss a defect. The tests may call the right code but make weak assertions, cover only one outcome, or never check a boundary condition. Mutation testing probes that gap by making small changes to production code and running the relevant tests against each changed version. If a test fails, the change is said to be killed. If all tests still pass, the mutation survives and deserves inspection.

Software Engineering 11 Sep 2026 9 min read

Fencing Tokens: Stop Stale Lock Holders from Writing

Fencing Tokens: Stop Stale Lock Holders from Writing A distributed lock can tell a client that it owns a resource for a limited period. That does not guarantee the client stops acting when the period ends. A process can pause for garbage collection, lose network access, become descheduled, or stall on an overloaded machine. During that pause, its lease can expire and another client can acquire the same lock. When the first process resumes, it may still believe it is entitled to write.

Software Engineering 11 Sep 2026 8 min read

Expand and Contract Database Changes for Safe Deployments

Expand and Contract Database Changes for Safe Deployments A database schema can change in milliseconds while an application fleet takes minutes or hours to converge on a new version. During that interval, old and new application instances may use the same database at the same time. That overlap turns an ordinary schema edit into a compatibility problem. Renaming a column in one migration, for example, can break old instances immediately even when the new application code is correct.

Software Engineering 11 Sep 2026 13 min read

Concurrency Limits: Bound In-Flight Work

Concurrency Limits: Bound In-Flight Work A service can receive traffic at an acceptable average rate and still collapse because too many operations overlap. The issue is not only how many requests arrive per second. It is also how many requests are active at the same time. A concurrency limit places a cap on active work. When all slots are occupied, additional work must wait, fail fast, or take another explicit path. This simple control can protect database connections, CPU-heavy routines, remote dependencies, worker capacity, and memory that grows with each active operation.

Software Engineering 10 Sep 2026 10 min read

Token Bucket Rate Limiting for Controlled Bursts

Token Bucket Rate Limiting for Controlled Bursts A service may handle 100 requests per second comfortably on average while still needing to accept a short burst of 300 requests after a client reconnects. A rigid per-second limit treats those situations as the same problem: once the current window is full, otherwise acceptable work is rejected. Token bucket rate limiting gives you a more useful control. It separates two decisions: how quickly permission to do work is replenished and how much permission may accumulate for a burst. Once you understand those two numbers, you can reason about the limiter without depending on a particular library or platform.

Software Engineering 10 Sep 2026 9 min read

Testing Asynchronous Behavior with Eventual Assertions

Testing Asynchronous Behavior with Eventual Assertions A test starts background work and then checks the result. On a fast machine the work finishes first and the test passes. Under CI load, the assertion runs a few milliseconds earlier and fails. Someone adds sleep(1 second). The failure disappears, but every successful run now pays a full second, and a sufficiently slow run can still fail. The problem is not that the test needs a longer delay. The test doesn’t know exactly when the result will become observable.

Software Engineering 10 Sep 2026 8 min read

Stale-While-Revalidate for Responsive Caches

Stale-While-Revalidate for Responsive Caches A cache entry expires just as a request arrives. The cached value is still only seconds old, but the request now has to wait while the application fetches a replacement from a slower dependency. If many entries expire during a busy period, cache refreshes can turn a normally fast read path into a burst of slow work. Stale-while-revalidate changes that trade-off. For data that can safely be slightly out of date, the application may return an expired cached value immediately while refreshing it separately for future requests. The reader gets predictable latency, and the cache still moves toward fresh data.

Software Engineering 10 Sep 2026 8 min read

Saga Pattern for Multi-Step Workflows

Saga Pattern for Multi-Step Workflows A workflow reserves inventory, charges a payment, and schedules delivery. Each step is handled by a different component with its own state. The inventory reservation succeeds, but payment fails. What should the system do with the reservation that already committed? A single database transaction can’t usually roll back work that has already been committed by independent components. The saga pattern handles this kind of workflow by treating it as a sequence of local transactions. When a later step fails, the workflow runs explicit compensating actions for earlier steps where business reversal is possible.

Software Engineering 10 Sep 2026 10 min read

Reconciliation Loops for Self-Healing Systems

Reconciliation Loops for Self-Healing Systems A one-shot operation works well when every step succeeds. Real systems are less cooperative. A process crashes after creating half its resources, an external API times out after accepting a request, an operator changes something manually, or a dependency becomes unavailable and recovers later. If correctness depends on one command completing perfectly, every interruption creates another recovery path to design and operate. A reconciliation loop uses a different model. Instead of saying, “perform these steps once,” the system repeatedly asks, “what should be true, what is true now, and what is the smallest safe action that moves reality toward the desired state?” That shift is useful for controllers, background jobs, provisioning systems, synchronizers, and any workflow where state can drift after the initial operation.

Software Engineering 10 Sep 2026 11 min read

Negative Caching for Repeated Failures

Negative Caching for Repeated Failures Caching usually brings successful results to mind: load a value once, keep it for a while, and avoid repeating expensive work. But repeated failures can be just as expensive as repeated successes. Suppose a service receives thousands of requests for an object that does not exist. If every request queries the same downstream system, the absence of that object becomes a source of load. The same pattern appears with invalid identifiers, unavailable optional resources, failed name lookups, and other outcomes that are expensive to rediscover but unlikely to change immediately.

Software Engineering 10 Sep 2026 9 min read

Differential Testing for Behavior-Preserving Changes

Differential Testing for Behavior-Preserving Changes Replacing working code is risky when the requirement is “change the implementation, not the behavior.” A rewritten parser may accept a different edge case. A faster pricing engine may round one value differently. A new library may return the same records in a different order. Ordinary tests help, but they only cover cases and assertions someone thought to write. Differential testing adds another source of evidence: run the old and new implementations on the same inputs, compare their observable results, and investigate differences.

Software Engineering 10 Sep 2026 10 min read

Designing Graceful Degradation for Partial Failures

Designing Graceful Degradation for Partial Failures A page needs product details, recommendations, reviews, and delivery estimates. The product service is healthy, but the recommendation service times out. Should the whole page fail? Sometimes yes. If the missing dependency is required to produce a correct result, failing the operation is the right behavior. But when the missing part is genuinely optional, turning one local failure into a complete outage throws away useful work.

Software Engineering 10 Sep 2026 10 min read

Design by Contract: Make Assumptions Explicit

Design by Contract: Make Assumptions Explicit A function often depends on rules that its type signature doesn’t fully express. A withdrawal amount must be positive. A completed operation must leave the balance consistent. An object may require its reserved quantity to stay between zero and the quantity on hand. When those rules live only in developers’ heads, failures appear far from their cause. Design by Contract gives the rules names and assigns responsibility for them. The useful mental model is simple: the caller promises to meet the operation’s entry conditions, and the operation promises a valid result while preserving the object’s valid state.