Skip to content

Archive

Software Engineering

462 articles
Software Engineering 23 Sep 2026 5 min read

Priority Inheritance Bounds Priority Inversion Around Mutexes

Priority Inheritance Bounds Priority Inversion Around Mutexes Priority scheduling does not guarantee that the highest-priority runnable task can always make progress. A high-priority task can block on a mutex held by a lower-priority task. If medium-priority work then preempts the owner, the high-priority task remains blocked even though the medium-priority work has no direct dependency on the mutex. This is priority inversion. The inversion starts with an ordinary dependency: the high-priority task needs a resource owned by a lower-priority task. The damaging part is interference from tasks between those priorities, which can delay the owner and extend the blocking interval.

Software Engineering 23 Sep 2026 6 min read

Lock Convoys Turn Short Critical Sections into Long Queues

Lock Convoys Turn Short Critical Sections into Long Queues A mutex can protect a tiny critical section and still become the center of a large latency problem. The code inside the lock may take only microseconds in normal operation, yet one delayed holder can allow several threads to accumulate behind it. Once that queue exists, the lock may remain continuously contended as ownership passes from one waiting thread to another.

Software Engineering 23 Sep 2026 6 min read

Hazard Pointers Delay Reclamation Until Readers Release References

Removing a node from a lock-free data structure does not make its memory immediately safe to reuse. Another thread may already hold the node’s address and may still dereference it. If the remover frees that allocation too early, an otherwise correct atomic update can be followed by a use-after-free. Hazard pointers separate logical removal from physical reclamation. A reader publishes the address it intends to access in a designated hazard slot. A remover can unlink a node and place it on a retired list, but reclamation waits until a scan confirms that no hazard slot protects that address.

Software Engineering 22 Sep 2026 6 min read

Write Skew Can Break Invariants Under Snapshot Isolation

Write Skew Can Break Invariants Under Snapshot Isolation Snapshot isolation gives each transaction a stable view of committed data and usually rejects concurrent updates to the same row. That combination removes many anomalies that appear under weaker isolation levels. It does not, however, make every application invariant serializable. Write skew is the important edge case. Two transactions read overlapping state, make decisions from the same valid snapshot, then update different rows. Because their write sets do not collide, both can commit. The combined result can violate a rule that neither transaction violated in its own snapshot.

Software Engineering 22 Sep 2026 6 min read

Transactional Outbox Closes the Database-to-Broker Commit Gap

Transactional Outbox Closes the Database-to-Broker Commit Gap A service often needs one operation to change database state and emit a message. An order may become confirmed while an OrderConfirmed event is sent to a broker. Those actions touch separate systems, so two ordinary writes cannot form one atomic commit unless both systems participate in a distributed transaction. The dangerous part is the interval between the writes. Commit the database first and the process can fail before publishing. Publish first and the database commit can fail afterward. Reversing the order moves the failure window; it does not remove it.

Software Engineering 22 Sep 2026 6 min read

Tombstones Prevent Deleted Data from Reappearing

Tombstones Prevent Deleted Data from Reappearing Deletion is not merely the absence of a value in a replicated store. Absence carries no information about whether a key was deliberately removed or whether a replica has simply never received it. When replicas can be temporarily disconnected, that distinction determines whether synchronization preserves a deletion or accidentally restores old data. A tombstone records the deletion as versioned state. Replicas can compare that marker with older values and keep the deletion when they reconcile. The marker can eventually be reclaimed, but only after the system has a defensible boundary beyond which an older value cannot return.

Software Engineering 22 Sep 2026 7 min read

Saga Compensation Is Not Transaction Rollback

Saga Compensation Is Not Transaction Rollback A multi-service operation can cross inventory, payments, shipping, and other independently committed systems. Once one service commits its step, a later failure cannot make that earlier commit disappear through an ordinary database rollback. A saga handles this boundary by pairing forward actions with explicit recovery actions. If a later step fails, the coordinator invokes compensations for earlier completed steps where the business process permits them.

Software Engineering 22 Sep 2026 7 min read

Request Sequence Guards Reject Late UI Responses

Request Sequence Guards Reject Late UI Responses Interactive interfaces often issue a new request before the previous one has finished. Search boxes, filters, route changes, autocomplete fields, and detail panels all create this pattern. Network completion order is not guaranteed to match the order in which the user changed state. That mismatch can produce a subtle race. Request A starts first, request B starts second, B completes first, and the interface renders B. If A then completes and its callback writes without checking context, the screen moves backward to data associated with the older intent.

Software Engineering 22 Sep 2026 7 min read

Request Coalescing Stops Cache Misses from Multiplying Backend Work

Request Coalescing Stops Cache Misses from Multiplying Backend Work A cache miss is usually cheap when one caller causes one backend lookup. The same miss can become expensive when many callers arrive for the same key at nearly the same time. Each caller observes the key as absent, each starts identical work, and the backend receives a burst precisely when the cache is providing no protection for that key. Request coalescing changes that concurrency pattern. The first caller becomes the leader for a key. Later callers join the same in-flight operation and wait for its result rather than starting equivalent work. Once the fill completes, the result can populate the cache and be returned to the waiting callers.

Software Engineering 22 Sep 2026 6 min read

Rendezvous Hashing Limits Key Movement During Membership Changes

Rendezvous Hashing Limits Key Movement During Membership Changes A partitioning rule has two jobs that can pull in different directions. It should spread keys across available nodes, and it should avoid moving most keys when that node set changes. A simple modulo rule handles the first job well for a stable cluster but performs poorly at the second. Rendezvous hashing, also called highest-random-weight hashing, assigns every key a deterministic score for every eligible node. The node with the highest score owns the key. Adding or removing a node changes only the comparisons involving that member, so keys with unaffected winners keep their placement.

Software Engineering 22 Sep 2026 8 min read

Optimistic Concurrency Rejects Stale Writes Before They Replace Newer State

Optimistic Concurrency Rejects Stale Writes Before They Replace Newer State A read-modify-write flow looks harmless when only one actor touches a record. A client reads state, changes part of it, then writes the result back. With concurrent actors, the interval between the read and the write becomes a race. Another writer can commit a newer value during that interval, and an unconditional update can erase it. Optimistic concurrency control puts a condition on the final write. The client carries a version derived from the state it read, and the storage layer accepts the mutation only if that version is still current. A mismatch becomes a conflict rather than a silent overwrite.

Software Engineering 22 Sep 2026 6 min read

Optimistic Concurrency Control Rejects Stale Writes

Optimistic Concurrency Control Rejects Stale Writes Two clients can read the same record, make different edits, and save seconds apart. If each update blindly replaces the stored value, the later write can erase the earlier one even though both requests succeeded. Optimistic concurrency control prevents that silent overwrite by attaching a condition to the write. The client records a version when it reads the data. Its update succeeds only if that version is still current. A changed version turns the write into a conflict instead of an unnoticed loss.

Software Engineering 22 Sep 2026 5 min read

Load Shedding Protects Services When Capacity Runs Out

Load Shedding Protects Services When Capacity Runs Out A service can receive more work than it can complete. The first visible symptom is often not an immediate error but a queue that grows while workers remain fully occupied. Requests spend longer waiting, deadlines expire, clients retry, and the extra retry traffic can deepen the overload. Load shedding places an explicit rejection point before that spiral consumes every available resource. The service admits work that fits its operating capacity and fails excess work quickly enough to preserve useful throughput for requests that can still complete.

Software Engineering 22 Sep 2026 6 min read

Idempotency Keys Turn Retries into One Logical Operation

Idempotency Keys Turn Retries into One Logical Operation A client can lose the response to a successful request. The server may commit a charge, create an order, or enqueue a job, then the connection can fail before the response reaches the caller. From the client’s view, success and failure are now ambiguous. Retrying is necessary for availability, but an ordinary retry can repeat the side effect. An idempotency key gives the client a way to say that several HTTP attempts represent one logical operation.

Software Engineering 22 Sep 2026 8 min read

Hedged Requests Trade Duplicate Work for Lower Tail Latency

Hedged Requests Trade Duplicate Work for Lower Tail Latency Most requests may finish quickly while a small fraction take much longer. A busy worker, a transient network queue, a cold cache entry, garbage collection, storage contention, or another local disturbance can stretch one attempt far beyond the median. At scale, those slow outliers become visible in p95, p99, and higher-percentile latency even when average service time looks healthy. A hedged request sends an additional attempt after the original has been outstanding for a chosen delay. Both attempts represent the same logical operation. The caller accepts the first valid result and cancels or ignores the remaining attempt.

Software Engineering 22 Sep 2026 6 min read

Fencing Tokens Stop Stale Lease Holders

Fencing Tokens Stop Stale Lease Holders A distributed lease gives one worker temporary permission to act as an owner. The lease eventually expires so another worker can take over after a crash or network failure. That solves availability, but expiry alone does not guarantee that the old worker has stopped. A process can pause long enough for its lease to expire, then resume with stale local state. A long garbage-collection pause, scheduler stall, suspended virtual machine, or delayed network path can create this condition. If the old worker writes after a replacement has taken ownership, two workers can affect the same resource even though the lease service never considered both leases valid at the same instant.

Software Engineering 22 Sep 2026 7 min read

Fencing Tokens Block Stale Lock Holders

Fencing Tokens Block Stale Lock Holders A distributed lock is often used to keep two workers from changing the same resource at once. The difficult case begins when lock ownership depends on a lease. A client can acquire the lease, pause long enough for it to expire, then resume after another client has acquired a new lease. From the old client’s point of view, execution simply continued. From the coordination service’s point of view, ownership already moved. If the protected storage system accepts both clients’ writes, the old holder can overwrite work performed by the current holder.

Software Engineering 22 Sep 2026 6 min read

Dead-Letter Queues Isolate Poison Messages Without Blocking Progress

Dead-Letter Queues Isolate Poison Messages Without Blocking Progress A message consumer usually treats failure as temporary at first. A database may be unavailable, a remote service may time out, or a worker may restart between receiving and acknowledging a message. Retrying is appropriate when another attempt has a reasonable chance of succeeding. Some messages fail for a different reason. Their payload is malformed, a referenced entity can never satisfy a required condition, or the consumer has a deterministic defect triggered by that input. Repeated delivery then consumes capacity without moving the message toward completion. A dead-letter queue gives that failure a separate destination after the normal retry policy is exhausted.

Software Engineering 22 Sep 2026 7 min read

Consistent Hashing Limits Key Movement During Topology Changes

Consistent Hashing Limits Key Movement During Topology Changes A distributed cache or partitioned service needs a rule that maps each key to a node. A simple rule such as hash(key) % N is attractive while the node count stays fixed. The trouble appears when N changes. Moving from four nodes to five changes the divisor for every key. Most remainders change, so a routine capacity adjustment can remap a large share of the dataset at once. For a cache, that can trigger a wave of misses. For stateful storage, it can create a large migration job.

Software Engineering 22 Sep 2026 7 min read

Circuit Breakers Stop Repeated Calls to Failing Dependencies

Circuit Breakers Stop Repeated Calls to Failing Dependencies A dependency that is already failing can consume more caller capacity than a healthy one. Requests wait for timeouts, retries add traffic, connection pools remain occupied, and worker slots stay tied to work that has little chance of completing. A circuit breaker places a stateful gate in front of that dependency so the caller can stop issuing calls after failure evidence reaches a configured limit.

Software Engineering 22 Sep 2026 6 min read

Bulkheads Keep One Saturated Dependency from Consuming Every Worker

Bulkheads Keep One Saturated Dependency from Consuming Every Worker A service can have healthy CPU, available memory, and responsive internal code while still becoming unavailable. One downstream dependency is enough to consume the service’s entire concurrency budget if calls to it become slow and every request is allowed to wait. The failure is not limited to the slow dependency. Shared worker pools, connection pools, semaphores, queues, and request slots turn local saturation into a service-wide outage. Bulkhead isolation limits that blast radius by reserving separate capacity for distinct workloads or dependencies.

Software Engineering 22 Sep 2026 5 min read

Bulkheads Isolate Concurrency Before One Dependency Consumes It All

Bulkheads Isolate Concurrency Before One Dependency Consumes It All A service can have enough CPU and memory yet stop making useful progress because its concurrency is exhausted. Threads, database connections, outbound sockets, worker slots, and in-flight request permits are finite. If one dependency becomes slow, calls to that dependency can occupy the entire shared pool. Bulkhead isolation divides that capacity before saturation occurs. Workloads that can fail independently receive separate concurrency budgets, so pressure in one path does not automatically consume every slot needed by another.

Software Engineering 22 Sep 2026 7 min read

Bounded Queues Turn Overload into an Explicit Admission Decision

Bounded Queues Turn Overload into an Explicit Admission Decision A queue absorbs short differences between arrival rate and service rate. That buffer is useful when a burst ends before workers fall far behind. The same mechanism becomes dangerous when arrivals remain faster than completions: every accepted item adds waiting time and consumes some combination of memory, descriptors, references, or durable storage. A bounded queue places a finite limit on that waiting population. Once the limit is reached, the system must make an admission decision instead of silently extending the backlog. Depending on the interface, that decision may block a producer, reject new work, shed selected work, or redirect it to another capacity domain.

Software Engineering 22 Sep 2026 6 min read

Backpressure Keeps Fast Producers from Overrunning Slow Consumers

Backpressure Keeps Fast Producers from Overrunning Slow Consumers A pipeline is stable only while work leaves each stage at roughly the rate it arrives over a useful time window. When a producer can submit work faster than a consumer can finish it, the difference has to accumulate somewhere. An unbounded queue makes that accumulation easy to miss. Requests continue to be accepted, the producer appears healthy, and the consumer keeps working. Meanwhile queued work consumes memory and ages before execution. Backpressure turns downstream saturation into an upstream signal before the backlog becomes the failure.