Skip to content

Archive

Software Engineering

462 articles
Software Engineering 20 Sep 2026 6 min read

Visibility Timeouts Turn Message Delivery into a Renewable Lease

A queue consumer often needs time to perform work before it can safely acknowledge a message. Removing the message at receive time would make a consumer crash capable of losing work. Keeping it immediately available would let several consumers process the same item at once. A visibility timeout occupies the middle ground. Receiving a message makes it temporarily unavailable to competing consumers. The consumer gets a bounded interval to finish and acknowledge it. If that interval expires first, the queue can expose the message for another delivery.

Software Engineering 20 Sep 2026 6 min read

Transactional Outbox Keeps Database State and Events Aligned

A service often needs one operation to change database state and emit an event. An order may move to paid while OrderPaid must reach a message broker. Those two writes cross different systems, so a normal database transaction cannot make both commits atomic. Writing the database first leaves a gap: the process can stop after commit but before publishing. Publishing first creates the opposite gap: consumers can observe an event for a database change that later fails.

Software Engineering 20 Sep 2026 7 min read

Tombstones Preserve Deletes Across Replicas Until Safe Garbage Collection

Tombstones Preserve Deletes Across Replicas Until Safe Garbage Collection Deleting a value from one copy of replicated data is not enough to delete it from the system. Another replica may be offline, delayed, or partitioned when the delete occurs. If the active replica simply removes the record, it also removes the evidence that a deletion happened. A stale replica can later return with an older value and make that value visible again.

Software Engineering 20 Sep 2026 6 min read

Token Buckets Separate Sustained Rate from Burst Capacity

A rate limit expressed only as “100 requests per second” leaves an important policy question open. Can a client send 100 requests at the first instant of each second, or must those requests be spread evenly? A token bucket makes that distinction explicit by separating sustained rate from burst capacity. The limiter maintains a balance of tokens up to a fixed capacity. Tokens arrive at a configured refill rate. An operation is admitted only when enough tokens are available, and admission deducts its cost from the balance. Idle time accumulates capacity for a later burst, but never beyond the bucket limit.

Software Engineering 20 Sep 2026 6 min read

Stale-While-Revalidate Keeps Cache Refresh off the Request Path

Stale-While-Revalidate Keeps Cache Refresh off the Request Path A cache entry does not become useless at the exact instant its freshness timer expires. For some data, a value that is a few seconds old is still preferable to making every caller wait for a backend refresh. Stale-while-revalidate uses that tolerance explicitly: the cache may serve an expired value for a bounded interval while a refresh runs separately. The policy changes refresh from a request-path requirement into background work for entries that remain acceptable while stale. It can reduce latency spikes around expiration, but only when the application can state how stale a response may become.

Software Engineering 20 Sep 2026 4 min read

Request Coalescing Collapses Concurrent Cache Misses

Request Coalescing Collapses Concurrent Cache Misses A cache miss can become expensive when many requests ask for the same key at nearly the same time. Without coordination, each caller can start an identical database query, remote call, or computation. The cache eventually fills, but the backend absorbs a burst precisely when the cached value is absent. Request coalescing changes the concurrency boundary. The first caller starts the load and publishes an in-flight entry for that key. Later callers join that entry instead of starting equivalent work. When the load finishes, its result is distributed to the waiting callers and the in-flight entry is removed.

Software Engineering 20 Sep 2026 6 min read

Rendezvous Hashing Keeps Key Placement Stable as Nodes Change

Rendezvous Hashing Keeps Key Placement Stable as Nodes Change Distributed systems often need a deterministic answer to a placement question: given a key and a current set of nodes, which node owns the key? A simple modulo rule such as hash(key) % N is compact, but changing N can move a large fraction of keys at once. Rendezvous hashing, also called highest-random-weight hashing, uses a different rule. For each key, it computes a deterministic score for every eligible node and selects the node with the highest score. Adding or removing a node changes placement only for keys whose ranking is affected by that membership change.

Software Engineering 20 Sep 2026 5 min read

Power of Two Choices Reduces Load Imbalance with Two Samples

Power of Two Choices Reduces Load Imbalance with Two Samples A load balancer that chooses one destination uniformly at random is cheap and decentralized, but random placement can produce uneven queues. At the other extreme, selecting the least loaded destination from the entire pool requires current load information for every candidate and can make the balancer itself expensive. The power-of-two-choices strategy sits between those designs. For each request, sample two eligible destinations, compare a load signal, and send the request to the better candidate. Two observations are enough to avoid many unlucky placements without requiring a global search.

Software Engineering 20 Sep 2026 7 min read

Load Shedding Protects Useful Work When Capacity Is Exhausted

A service can be healthy at 2,000 requests per second and collapse at 2,400. The extra 400 requests do not merely wait their turn. They may occupy connection slots, queue entries, memory, worker threads, database sessions, and retry budgets while useful throughput falls. Load shedding places an explicit admission decision before a scarce resource is fully consumed. When the system cannot serve all incoming work within its operating envelope, it rejects selected requests early instead of allowing every request to compete until they all become slow.

Software Engineering 20 Sep 2026 6 min read

Lease Renewal Needs a Safety Margin Before Expiry

Lease Renewal Needs a Safety Margin Before Expiry A lease grants a holder temporary authority until a recorded expiry. Keeping that authority requires renewal before the deadline. Scheduling renewal at the deadline itself leaves no room for network delay, scheduler pauses, storage latency, or a transient retry. A safer design attempts renewal earlier. The interval between the planned renewal and expiry is a safety margin: time reserved for ordinary uncertainty before the lease is treated as lost.

Software Engineering 20 Sep 2026 7 min read

Idempotency Keys Make Retried Writes Safe to Repeat

Idempotency Keys Make Retried Writes Safe to Repeat A client can lose the response to a successful write. The connection may close after the server commits a payment, creates an order, or schedules a job but before the response reaches the caller. From the client’s perspective, failure and success can look identical. Retrying blindly is dangerous for operations with non-idempotent effects. Sending the same POST twice may create two resources or charge twice. Refusing to retry leaves the caller with an ambiguous outcome.

Software Engineering 20 Sep 2026 8 min read

Idempotency Keys Make Retried Commands Safe

A client can lose the response to a command even when the server completed the operation. The connection may close after a payment is recorded, a job is created, or an order is accepted. From the client’s point of view, timeout does not reveal whether the command failed before execution or succeeded before the response disappeared. Retrying is necessary for availability, but repeating a state-changing command can duplicate the side effect. An idempotency key gives all attempts for one logical command the same identity. The server stores the outcome associated with that identity and reuses it when the same command arrives again.

Software Engineering 20 Sep 2026 5 min read

Hedged Requests Trade Extra Work for Lower Tail Latency

A service can have a healthy median latency and still produce occasional requests that take far longer than the rest. Queueing, a slow replica, connection setup, garbage collection, storage stalls, or transient network delay can leave one attempt behind while equivalent capacity elsewhere remains available. A hedged request limits exposure to that single slow path. The client starts one attempt normally. If it is still pending after a configured delay, the client may start a second equivalent attempt. The first acceptable result is used, and the remaining attempt is cancelled when cancellation is supported.

Software Engineering 20 Sep 2026 6 min read

Hedged Requests Cut Tail Latency at a Controlled Cost

Hedged Requests Cut Tail Latency at a Controlled Cost A service can have acceptable median latency and still produce a small set of very slow responses. Queueing, runtime pauses, storage contention, packet loss, or a temporarily busy replica can push individual requests far beyond the common case. Hedged requests address that tail by sending a second copy after the first request has been outstanding for a chosen delay. The caller accepts the first valid response and cancels the remaining attempt. The technique trades a bounded amount of extra work for a chance to escape an unusually slow execution path.

Software Engineering 20 Sep 2026 7 min read

Fencing Tokens Block Stale Lock Holders at the Resource

Fencing Tokens Block Stale Lock Holders at the Resource A distributed lease can expire while its holder is still running. A process may pause for garbage collection, lose contact with the coordinator, stall under scheduler pressure, or resume after a machine suspension. The lock service can correctly grant the lease to another worker while the old worker still has unfinished work. That gap matters when both workers can reach the protected resource. A lease controls ownership in the coordinator; it does not automatically revoke a delayed database connection, storage request, or RPC that was prepared by the previous holder.

Software Engineering 20 Sep 2026 7 min read

Deadline Propagation Stops Work After Callers Give Up

Deadline Propagation Stops Work After Callers Give Up A timeout at the edge does not automatically stop work deeper in a system. A client may abandon a request after two seconds while an API server continues waiting on another service, which may still be running a database query. The response has lost its consumer, yet CPU time, connections, memory, queue positions, and downstream capacity can remain occupied. Deadline propagation carries the caller’s time budget across those boundaries. Each component receives an absolute deadline or an equivalent remaining budget, refuses work that cannot start in time, and cancels operations when the budget expires. The goal is not merely faster failure. It is to keep useless work from surviving longer than the request that justified it.

Software Engineering 20 Sep 2026 6 min read

Deadline Propagation Stops Expired Requests from Consuming Downstream Capacity

A timeout placed only at the outer edge of a request does not automatically limit the work started deeper in the call graph. The client may stop waiting after 800 milliseconds while an internal service continues a database query, a remote call, or a queued task for several more seconds. The response is already useless to that client, yet the system is still spending capacity on it. Deadline propagation carries the request’s time boundary with the work. Each component can compare that boundary with its current clock, reserve time for its own processing, and refuse or cancel work that no longer fits. The result is not merely faster failure. It is a tighter relationship between useful work and resource consumption.

Software Engineering 20 Sep 2026 6 min read

Consistent Hashing Limits Key Movement When Nodes Change

Partitioning by hash(key) % N is simple when the node count stays fixed. The arithmetic becomes disruptive when N changes. Moving from four nodes to five changes the divisor, so many keys select a different remainder even though only one node joined. Consistent hashing changes the mapping. Keys and nodes are placed in the same circular hash space. A key belongs to the first node encountered in a chosen direction around the ring. Adding or removing a node changes ownership only for ranges adjacent to that membership change.

Software Engineering 20 Sep 2026 4 min read

Circuit Breakers Limit Repeated Calls to Failing Dependencies

Circuit Breakers Limit Repeated Calls to Failing Dependencies A remote dependency can fail in a way that is both slow and expensive. Requests wait for timeouts, workers remain occupied, retries add more traffic, and a local service can lose capacity even when its own code is healthy. A circuit breaker places a stateful decision in front of that call path. While the dependency behaves acceptably, calls pass through. After the configured failure condition is reached, the breaker opens and rejects new calls locally for a bounded period. Later, it admits a small number of probes before deciding whether normal traffic can resume.

Software Engineering 20 Sep 2026 7 min read

Bulkheads Isolate Resource Pools Before Failures Spread

A service can remain healthy at the process level while becoming useless because one workload has consumed every scarce execution resource. A slow dependency can occupy all outbound connections. A noisy tenant can fill every worker slot. A background job can take the same semaphore permits needed by interactive requests. Bulkhead isolation limits that coupling. Instead of letting unrelated work compete for one undifferentiated pool, the system partitions selected resources and gives each class of work a bounded share. Saturation then has a smaller blast radius.

Software Engineering 20 Sep 2026 5 min read

Bulkheads Isolate Concurrency Across Dependencies

Bulkheads Isolate Concurrency Across Dependencies A service can have plenty of CPU and still become unavailable because one dependency stops completing work. Requests waiting on a slow database, remote API, or storage service retain execution slots, connections, memory, and queue positions. If unrelated operations share the same finite pool, one saturated path can consume capacity needed by healthy paths. Bulkhead isolation divides that shared concurrency into explicit budgets. Calls to one dependency or workload class use a bounded pool that other classes cannot exhaust. The pattern does not repair a failing dependency. It limits the amount of local capacity that failure can occupy.

Software Engineering 20 Sep 2026 7 min read

Bounded Queues Turn Overload into Explicit Rejection

Bounded Queues Turn Overload into Explicit Rejection A queue can absorb short bursts when requests arrive faster than workers can finish them. That buffer is useful only while it remains a buffer. If producers can keep adding work without a fixed limit, sustained overload turns the queue into an expanding inventory of requests that may wait long after their results are useful. A bounded queue changes the failure mode. It accepts waiting work up to a deliberate capacity, then refuses additional admission until space becomes available. The service still experiences overload, but the overload appears as an explicit control decision rather than unbounded growth in memory and waiting time.

Software Engineering 20 Sep 2026 6 min read

Backpressure Keeps Producer Speed Tied to Consumer Capacity

A fast producer and a slower consumer can coexist safely only while the gap between their rates remains bounded. If incoming work arrives faster than it can be completed for long enough, buffering does not remove overload. It stores the difference. Backpressure makes that capacity mismatch part of the protocol between components. Instead of accepting work indefinitely, a saturated stage causes upstream code to slow down, wait for capacity, reduce demand, or reject work according to an explicit policy.

Software Engineering 20 Sep 2026 7 min read

Adaptive Concurrency Limits Follow Service Capacity

Adaptive Concurrency Limits Follow Service Capacity A service can become slower before it becomes unavailable. As in-flight work rises, CPU queues grow, connection pools fill, lock contention increases, and downstream calls accumulate. A fixed concurrency ceiling can protect the service, but one number rarely fits every operating condition. Capacity shifts with request mix, cache hit rate, dependency latency, deployment shape, and resource pressure. Adaptive concurrency control treats the admission limit as a value that can move. The controller observes recent service behavior, raises the limit while additional concurrency remains productive, and reduces it when latency indicates growing queues or saturation. The goal is not maximum concurrency. It is enough parallel work to use available capacity without allowing queues to dominate response time.