Skip to content

Topic archive

Software Engineering

Software Engineering covers developer tooling, architecture, maintainability, engineering workflows, protocols, and practices that improve how software is designed and built.

460 articles
Software Engineering 21 Sep 2026 8 min read

Write Skew Breaks Invariants Under Snapshot Isolation

Write Skew Breaks Invariants Under Snapshot Isolation Snapshot isolation gives each transaction a stable view of committed data. That property removes many anomalies caused by values changing midway through a transaction. It does not, by itself, make every concurrent execution equivalent to some serial order. Write skew is a compact example of the gap. Two transactions read the same valid state, make decisions from that state, then write different rows. Because their write sets do not overlap, both commits can succeed even though the combined result violates a rule that each transaction preserved in isolation.

Software Engineering 21 Sep 2026 5 min read

Write Skew Breaks Cross-Row Invariants Under Snapshot Isolation

Write Skew Breaks Cross-Row Invariants Under Snapshot Isolation Snapshot isolation gives each transaction a stable database view and commonly prevents concurrent transactions from committing conflicting writes to the same row. That is a strong concurrency property, but it does not make every application invariant serializable. Write skew appears when two transactions read overlapping state, make decisions from the same valid snapshot, then write different records. Since their write sets do not collide, both commits may succeed. The combined state can violate a rule that each transaction checked before writing.

Software Engineering 21 Sep 2026 7 min read

Version Vectors Separate Causality from Concurrency

Version Vectors Separate Causality from Concurrency Replicated data can receive writes at several nodes while communication between those nodes is delayed. When two versions meet later, a store has to decide whether one descends from the other or whether both were created independently. A wall-clock timestamp gives a total-looking order, but clock order is not causal order. Two replicas can write during a partition, and whichever timestamp happens to be larger does not make that write a descendant of the other.

Software Engineering 21 Sep 2026 4 min read

Version Vectors Distinguish Concurrent Updates from Causal Successors

Version Vectors Distinguish Concurrent Updates from Causal Successors Replicated data can receive writes at different nodes while communication between those nodes is delayed. When versions later meet, a scalar revision number can say that two values differ, but it cannot always say whether one descends from the other or both were produced independently. A version vector records progress per replica. Comparing those counters provides a partial order: one version can dominate another, the vectors can be equal, or neither can dominate. The last case identifies concurrent histories that require an explicit reconciliation rule.

Software Engineering 21 Sep 2026 5 min read

Version Columns Turn Lost Updates into Detectable Conflicts

Version Columns Turn Lost Updates into Detectable Conflicts A read-modify-write flow can overwrite another committed change even when every individual database statement succeeds. Two clients read the same row, compute different replacements, then write in sequence. Without a condition tying each write to the state it read, the later write can silently erase the earlier one. A version column makes that dependency explicit. The client reads both data and version, then updates only if the stored version is still the one it observed. A changed version turns the race into a failed conditional update instead of a lost update.

Software Engineering 21 Sep 2026 8 min read

Transactional Outbox Closes the Dual-Write Gap

Transactional Outbox Closes the Dual-Write Gap A service often needs one request to change database state and publish an event. The two operations may look adjacent in application code, but they cross different durability boundaries. A database commit can succeed while a broker publish fails, or the publish can succeed before the database transaction rolls back. That split creates a dual-write problem. No ordering of two independent writes can make them atomic by itself.

Software Engineering 21 Sep 2026 5 min read

Transactional Outbox Closes the Database-Broker Commit Gap

Transactional Outbox Closes the Database-Broker Commit Gap A service often needs one request to change database state and publish a message. Those actions may look adjacent in application code, but they cross two independent commit boundaries. If the database and broker do not share a transaction protocol, no ordering of two ordinary writes can make them atomic. Consider an order service that stores an accepted order and emits OrderCreated. Publishing after the database commit leaves a crash window before the broker call. Publishing first creates the opposite window: consumers can receive an event for state that later fails to commit.

Software Engineering 21 Sep 2026 6 min read

Token Buckets Preserve Burst Capacity Without Removing Rate Bounds

Token Buckets Preserve Burst Capacity Without Removing Rate Bounds A fixed requests-per-second ceiling treats a brief spike and a sustained flood as the same event. That can be too rigid for services whose callers naturally arrive in clusters. A token bucket separates two constraints: the long-run admission rate and the amount of burst traffic the service is willing to absorb. The model has two parameters. The bucket capacity B is the maximum number of tokens that can accumulate. The refill rate r adds tokens per unit of time, up to B. A request consumes tokens according to its configured cost. If enough tokens are present, the request proceeds; otherwise it is rejected, delayed, or handled by another explicit policy.

Software Engineering 21 Sep 2026 6 min read

Request Coalescing Collapses Cache-Miss Bursts

Request Coalescing Collapses Cache-Miss Bursts A cache miss is usually cheap when one caller triggers one backend read. The same miss can become expensive when hundreds of callers arrive for the same key before the first fill completes. Each caller sees an empty cache and starts equivalent work, multiplying load precisely when the cached value is unavailable. Request coalescing places a small concurrency boundary around that fill. The first caller starts the backend operation. Later callers for the same key join the in-flight operation instead of starting another one. When it completes, the result can populate the cache and be delivered to the waiting callers.

Software Engineering 21 Sep 2026 7 min read

Load Shedding Protects Useful Work Under Saturation

Load Shedding Protects Useful Work Under Saturation A service has a finite amount of work it can complete per unit of time. When offered load rises past that capacity, accepting every request does not create more capacity. It creates more waiting, consumes memory and connection slots, extends deadlines, and can leave expensive work running after callers have already given up. Load shedding makes admission explicit. Work that the service cannot process within its operating envelope is rejected early so admitted work retains a realistic chance of completing.

Software Engineering 21 Sep 2026 7 min read

Idempotency Keys Make Retried Mutations Safe

Idempotency Keys Make Retried Mutations Safe A client can lose the result of a successful mutation without losing the mutation itself. The server may commit a payment, reservation, or job submission and then drop the connection before the response reaches the caller. From the client’s perspective, a timeout leaves two plausible states: the operation failed before commit, or it committed and only the response was lost. Blindly retrying a non-idempotent mutation can apply the effect twice. Refusing every retry leaves the caller unable to recover from an ambiguous outcome. An idempotency key gives both sides a stable identity for one logical operation, so a repeated attempt can reuse the result of the first accepted attempt instead of creating another effect.

Software Engineering 21 Sep 2026 6 min read

Idempotency Keys Bound Duplicate Mutations Across Retries

Idempotency Keys Bound Duplicate Mutations Across Retries A client can lose the response to a successful mutation. The server may commit a payment, reservation, or job submission and then lose the connection before the response reaches the caller. From the client side, timeout does not reveal whether the mutation failed before commit or succeeded before the response disappeared. A retry is necessary for availability, but a blind retry can repeat the side effect. An idempotency key gives both attempts a stable identity so the server can treat them as one logical operation.

Software Engineering 21 Sep 2026 7 min read

Hedged Requests Cut Tail Latency with Controlled Duplication

Hedged Requests Cut Tail Latency with Controlled Duplication A service can have a healthy median latency while a small fraction of requests take far longer. Queueing, garbage collection, storage stalls, packet loss, noisy neighbors, or uneven replica load can all stretch the slow end of the distribution. For a request that fans out to several dependencies, one slow branch can dominate the entire response. Hedged requests reduce that exposure by starting a second copy after a short delay. The copies target independent execution paths when possible, and the first valid response wins.

Software Engineering 21 Sep 2026 6 min read

Hedged Requests Cut Tail Latency at a Capacity Cost

Hedged Requests Cut Tail Latency at a Capacity Cost A service can have acceptable median latency while a small fraction of requests take much longer. Queueing, a cold cache, runtime pauses, transient packet loss, or a slow storage operation can leave one attempt far behind the normal path. At sufficient fan-out, those rare delays become common at the aggregate request boundary. A hedged request starts a second equivalent attempt after the first has remained incomplete for a selected delay. The caller accepts the first useful result and cancels or discards the other attempt.

Software Engineering 21 Sep 2026 6 min read

Fencing Tokens Stop Stale Lease Holders from Writing

Fencing Tokens Stop Stale Lease Holders from Writing A distributed lease can decide which client currently owns a resource, but lease expiry does not instantly stop the previous holder. A process can pause, lose network access, or stall long enough for its lease to expire. Another client then acquires the lease. If the old process resumes and still has access to the protected storage or service, both clients can issue writes.

Software Engineering 21 Sep 2026 5 min read

Deadline Propagation Preserves Timeout Budgets Across RPC Hops

Deadline Propagation Preserves Timeout Budgets Across RPC Hops A timeout that restarts at every service boundary can turn a short caller budget into a much longer chain of work. A client may allow 800 milliseconds, service A may spend 500 milliseconds locally, then call service B with a fresh 800-millisecond timeout. B can continue working long after the client has stopped waiting. Deadline propagation keeps one end time attached to the request. Each hop derives its remaining budget from that deadline and refuses to start work that cannot fit within it. The result is not faster execution by itself. It is bounded execution that respects the time contract established upstream.

Software Engineering 21 Sep 2026 6 min read

Deadline Propagation Preserves Request Time Budgets

Deadline Propagation Preserves Request Time Budgets A request can cross several services before producing a response. Each hop may have its own queue, network call, retry policy, and local timeout. If those limits are chosen independently, the total path can run far longer than the caller is prepared to wait. Deadline propagation gives the path one temporal boundary. The initiating caller supplies an absolute deadline, or a time budget that is converted into one. Each downstream component uses the remaining interval rather than starting a fresh timeout from zero.

Software Engineering 21 Sep 2026 6 min read

Consistent Hashing Limits Key Movement During Membership Changes

Consistent Hashing Limits Key Movement During Membership Changes A simple hash partition often looks sufficient: owner = hash(key) % node_count With four nodes, every key maps to one of four remainders. The problem appears when the membership changes. Moving from four nodes to five changes the modulus, so a large fraction of keys select a different owner even though only one node was added.

Software Engineering 21 Sep 2026 6 min read

Circuit Breakers Limit Cascading Failure

Circuit Breakers Limit Cascading Failure A slow or failing dependency can consume more than its own capacity. Callers wait, retry, hold sockets, occupy worker slots, and retain memory while requests remain unresolved. As pressure spreads upstream, a local fault can become a service-wide saturation event. A circuit breaker places a stateful gate around calls to that dependency. It observes outcomes, opens when the configured failure policy is met, rejects calls for a period, then permits a small number of probes. Successful probes can return the breaker to normal traffic; failed probes send it back to the open state.

Software Engineering 21 Sep 2026 7 min read

Circuit Breakers Bound Calls to an Unhealthy Dependency

Circuit Breakers Bound Calls to an Unhealthy Dependency A remote dependency can fail in a way that is both persistent and expensive. Connections time out, worker slots remain occupied, request queues grow, and retries add more traffic to a service that is already unable to respond. A caller that keeps issuing the same class of request can turn one dependency failure into pressure across its own process. A circuit breaker places a stateful admission decision in front of those calls. While the dependency behaves within policy, calls pass through. After enough qualifying failures, the breaker opens and rejects new calls locally for a period. Recovery is tested with limited traffic rather than a full return to normal load.

Software Engineering 21 Sep 2026 7 min read

Bulkheads Isolate Concurrency Across Failure Domains

Bulkheads Isolate Concurrency Across Failure Domains A service can remain reachable while its useful capacity disappears. A slow dependency holds requests open, those requests occupy workers or connection slots, and unrelated traffic waits behind work that cannot finish promptly. The fault began in one path, but a shared resource pool lets it consume capacity needed by every path. Bulkhead isolation partitions that finite capacity. Calls associated with one failure domain receive a bounded share rather than competing without separation for the entire pool. When one partition fills, admission fails or waits within that partition while capacity assigned to other work remains available.

Software Engineering 21 Sep 2026 6 min read

Bounded Queues Turn Overload into Explicit Backpressure

Bounded Queues Turn Overload into Explicit Backpressure A queue can absorb a short mismatch between arrival rate and processing rate. That buffering is useful when bursts are temporary. It becomes dangerous when the queue has no meaningful bound: sustained overload no longer appears as an admission failure, but as a growing backlog, rising memory use, and requests that finish long after their latency budget has expired. A bounded queue changes the contract. Once capacity is exhausted, the producer must wait, reject, shed, or route work elsewhere. The overload is no longer hidden inside an expanding buffer.

Software Engineering 21 Sep 2026 6 min read

Backpressure Bounds Work When Consumers Fall Behind

Backpressure Bounds Work When Consumers Fall Behind A fast producer and a slower consumer can coexist for a short burst if a buffer absorbs the difference. The same arrangement becomes unstable when the rate mismatch persists. Pending work accumulates, memory rises, latency stretches, and items may expire before a consumer reaches them. Backpressure changes the contract between both sides. Instead of accepting work without regard to downstream state, the system exposes limited capacity to the producer. When that capacity is exhausted, production pauses, admission is rejected, or another explicit overload policy takes effect.

Software Engineering 21 Sep 2026 7 min read

Adaptive Concurrency Limits Track Available Service Capacity

Adaptive Concurrency Limits Track Available Service Capacity A fixed concurrency ceiling is easy to operate when service capacity is stable. Real systems rarely stay in one operating regime. Database contention, cache hit rate, request mix, downstream latency, CPU availability, and deployment changes can all move the amount of work a service can sustain at once. Adaptive concurrency control treats the in-flight limit as a control variable. The limiter admits work up to a current ceiling, observes service behavior, then adjusts that ceiling. The aim is not maximum concurrency. It is enough concurrency to use available capacity without allowing queues to grow far beyond the useful operating region.