Skip to content

Archive

Software Engineering

462 articles
Software Engineering 12 Sep 2026 10 min read

Prevent Write Skew with Serializable Transactions

Database transactions make many state changes easier to reason about, but transaction boundaries alone do not guarantee that every business invariant survives concurrency. A particularly subtle failure is write skew: two transactions read overlapping state, update different rows, and both commit even though their combined result violates a rule. This anomaly matters because each transaction can look correct in isolation. The defect appears only when valid decisions are made from snapshots that become incompatible once both writes are accepted.

Software Engineering 12 Sep 2026 10 min read

Phantom Rows and the Limits of Row-Level Locking

A transaction can lock every row it reads and still leave a business rule exposed. The gap appears when the rule is about a set described by a predicate, not only the rows that currently satisfy it. Suppose an application limits a small allocation group to four active reservations. A transaction queries the active rows, sees three, and decides that one more reservation is valid. If another transaction inserts a new matching row before the first transaction commits, both decisions may have been based on a set that no longer represents the committed state.

Software Engineering 12 Sep 2026 9 min read

Optimistic Concurrency with Version Columns

A row can be read correctly, modified correctly, and still be written incorrectly. The problem appears when another transaction changes the same logical record between the read and the write. A plain UPDATE often has no memory of the state on which the new values were based, so the later writer can replace an earlier change without detecting the race. A version column turns that hidden assumption into a predicate. The update says, in effect, that the write is valid only while the row remains at the version that the caller observed. The database then evaluates the state check and the mutation as one atomic statement.

Software Engineering 12 Sep 2026 11 min read

Load Shedding: Reject Work Before Overload Spreads

Load Shedding: Reject Work Before Overload Spreads A service has finite capacity. When offered work exceeds that capacity for long enough, accepting every request can make the service less useful rather than more useful. Queues grow, deadlines expire, memory pressure rises, dependencies receive more traffic, and successful throughput can fall. Load shedding is the deliberate rejection of work that the system cannot serve within an acceptable budget. The goal is not to maximize the number of requests admitted. The goal is to preserve useful service during overload.

Software Engineering 12 Sep 2026 11 min read

Idempotent Consumers: Handle Duplicate Messages Safely

Idempotent Consumers: Handle Duplicate Messages Safely A message broker can deliver the same message more than once. A worker may finish its database update and crash before acknowledging the message. The broker sees no acknowledgement, so it sends the message again. From the broker’s perspective, redelivery is the safe choice. From the application’s perspective, the second delivery can repeat a business effect. That gap matters whenever an effect must happen once per logical message. Charging an account twice, granting stock twice, incrementing a counter twice, or sending the same fulfillment request twice can turn a routine retry into corrupted state.

Software Engineering 12 Sep 2026 9 min read

Fencing Tokens: Block Stale Lease Holders

Fencing Tokens: Block Stale Lease Holders A distributed lease can grant one process temporary permission to act, but expiration alone cannot stop that process from acting after its lease has ended. A long pause, network delay, overloaded runtime, or suspended virtual machine can leave an old holder unaware that another process has already acquired the lease. This creates a subtle safety gap. Two processes can both believe they are entitled to modify the same resource, even when the lease service itself grants ownership correctly.

Software Engineering 12 Sep 2026 8 min read

Fencing Tokens for Expiring Distributed Leases

A lease can expire while its holder is still running. That single property separates a distributed lease from an ordinary in-process mutex. The coordinator may grant ownership to another client after a deadline, yet the former holder can resume after a long pause and continue issuing operations based on authority it no longer has. The coordinator has done its job: it stopped treating the old client as the current holder. The shared resource has a different problem. Unless operations carry evidence of ownership order, the resource may have no basis for distinguishing a current holder from a stale one.

Software Engineering 12 Sep 2026 7 min read

Expand and Contract at Database Schema Boundaries

A database column can be structurally valid and still be incompatible with the application processes using it. Renaming customer_name to display_name, for example, is trivial as a data-definition operation on many databases. The harder boundary appears when one application process still issues queries against the old name while another process already expects the new one. That overlap is common whenever application replacement is not atomic. Rolling deployments, multiple service instances, delayed workers, and independent consumers can leave more than one application version active at the same time. A schema migration then has two audiences: the database engine and every executable version that can reach the database during the transition.

Software Engineering 12 Sep 2026 7 min read

Deadline Propagation as a Request Boundary

Deadline Propagation as a Request Boundary A service can return after its caller has stopped waiting. The computation may still consume a connection, hold a concurrency slot, execute a database query, or start another remote call. A local timeout limits how long one caller waits; it does not, by itself, bound the lifetime of work already sent deeper into the system. An end-to-end deadline changes that boundary. Instead of giving each operation an independent duration, the request carries a point in time after which its result is no longer useful to the initiating operation. Each component can derive its remaining budget from that same boundary.

Software Engineering 12 Sep 2026 9 min read

Consistent Hashing: Limit Key Movement as Nodes Change

Distributed systems often need a deterministic answer to a simple question: given a key, which node should own it? A cache cluster may route each object key to one server. A storage service may assign each partition to a shard. A worker pool may send all events for the same account to the same processor. The routing rule must be stable enough that clients agree, yet flexible enough to handle nodes joining and leaving.

Software Engineering 12 Sep 2026 8 min read

Conditional HTTP Writes with Entity Tags

Conditional HTTP Writes with Entity Tags A client reads a resource, edits its local copy, and sends a replacement several seconds later. During that interval another client may have committed a different replacement. A plain PUT has no statement about the representation on which the edit was based, so the server can accept a request whose starting state is already obsolete. HTTP conditional requests can carry that missing premise. A response entity tag identifies a selected representation, and If-Match makes a later request conditional on a current representation matching one of the supplied tags. For state-changing methods, that turns representation identity into an explicit concurrency boundary.

Software Engineering 12 Sep 2026 7 min read

Compensation Is Not Rollback Across Service Boundaries

Compensation Is Not Rollback Across Service Boundaries A local database rollback can erase uncommitted writes before other transactions are allowed to depend on them. A compensating operation has a different shape. It runs after an earlier operation has committed, often after that result has become visible to other components. That distinction changes the consistency model. Compensation does not restore a distributed system to a state in which the original action never occurred. It adds another state transition whose domain meaning offsets some consequence of the first one.

Software Engineering 12 Sep 2026 9 min read

Circuit Breakers as Admission Control for Failing Dependencies

Circuit Breakers as Admission Control for Failing Dependencies A remote call that has little chance of succeeding still consumes something: a connection slot, a worker, a deadline budget, memory for request state, or capacity in the dependency itself. When repeated failures indicate that a downstream service is currently unable to serve useful work, continuing to admit every call can preserve the very pressure that callers need to escape. A circuit breaker changes that admission decision. Instead of treating each call as independent, it retains a small amount of state about recent outcomes. That state can temporarily reject new calls before network I/O begins, then permit controlled probes after a recovery interval.

Software Engineering 12 Sep 2026 10 min read

Cache Stampede Control with Early Refresh

Cache Stampede Control with Early Refresh A cache can remove enormous amounts of repeated work, yet a popular entry creates a sharp risk at expiry. If ten thousand requests depend on the same key and that key expires, many requests can discover the miss at nearly the same moment. Each request may then start the same database query, computation, or remote call. This event is commonly called a cache stampede. The cache works well during the entry lifetime, then abruptly stops protecting the dependency exactly when demand is high.

Software Engineering 12 Sep 2026 11 min read

Bulkhead Isolation: Contain Failures with Separate Capacity Pools

Bulkhead Isolation: Contain Failures with Separate Capacity Pools A service can have enough total capacity and still become unavailable because one dependency consumes all of it. Imagine an API that calls a payment service and a recommendation service. Both outbound calls use the same worker pool. Recommendations become slow during a traffic spike. Their requests occupy every worker while waiting for responses. Payment requests now have no worker available, even though the payment service itself is healthy.

Software Engineering 12 Sep 2026 10 min read

Backpressure: Match Producer Speed to Consumer Capacity

Backpressure: Match Producer Speed to Consumer Capacity A pipeline is stable only when work enters at a rate its downstream stages can sustain. That sounds obvious, yet many systems let producers run at full speed until a queue fills, memory grows, latency explodes, or a downstream service starts rejecting requests. The visible failure appears late. The actual mismatch began earlier: one stage could create work faster than the next stage could finish it.

Software Engineering 11 Sep 2026 8 min read

Using a Facade to Reduce Dependency Surface

Using a Facade to Reduce Dependency Surface A feature starts with one call into a library. A few months later, every caller knows which three components to create, which methods must run first, which defaults belong together, and which low-level errors need translation. The subsystem still works, but its internal structure has leaked into the rest of the codebase. A facade is a deliberately simpler interface placed in front of a more complicated subsystem. Callers depend on the operations they actually need instead of coordinating the subsystem themselves. This article explains how to recognize that design problem, build a focused facade, and decide when the extra layer is useful rather than ceremonial.

Software Engineering 11 Sep 2026 9 min read

Tell, Don’t Ask: Keep Decisions with the Object That Owns the Data

Tell, Don’t Ask: Keep Decisions with the Object That Owns the Data A class can expose perfectly reasonable getters and still make a codebase harder to change. The trouble appears when callers fetch several values, interpret them, make a business decision, and then tell the object how to update itself. The data lives in one place, but the rule that gives the data meaning lives somewhere else. Tell, Don’t Ask is a design guideline for reducing that split. Instead of asking an object for internal state so a caller can decide what should happen, prefer telling the object the meaningful operation you want performed. The object can then apply the rules that belong with its state.

Software Engineering 11 Sep 2026 9 min read

Stable Dependencies Principle: Point Toward Stability

Stable Dependencies Principle: Point Toward Stability A dependency graph is not only a map of which component calls which other component. Its direction also determines which parts of a system can change independently. Consider a reporting component used by ten other components. Many callers depend on it, so changing its public contract can require coordinated work across the codebase. That reporting component has become relatively stable: not necessarily because its code rarely changes, but because many other components constrain how freely its contract can change.

Software Engineering 11 Sep 2026 8 min read

Shotgun Surgery: When One Change Touches Many Files

Shotgun Surgery: When One Change Touches Many Files A pricing rule changes from “free shipping above 50” to “free shipping above 75.” The rule is simple, but the pull request touches checkout, order previews, invoice generation, customer notifications, and tests in several modules. Miss one place and customers see contradictory behavior. That pattern is called shotgun surgery: one logical change requires many small edits scattered through the codebase. The problem isn’t the number of files by itself. The problem is that one decision has many owners.

Software Engineering 11 Sep 2026 10 min read

Shotgun Surgery: Reduce Scattered Change

Shotgun Surgery: Reduce Scattered Change A small requirement can produce a surprisingly large patch. Adding one order state means editing validation, formatting, notification, audit, and reporting code in different modules. Each edit may be simple, yet missing one can leave the system inconsistent. This recurring shape is often called shotgun surgery: one conceptual change forces small edits across many places. The practical problem isn’t the number of files by itself. It is that a single responsibility has been scattered, so developers must reconstruct the full change surface each time that responsibility evolves.

Software Engineering 11 Sep 2026 7 min read

Replacing Type Branches with Polymorphism

Replacing Type Branches with Polymorphism A type-based switch can be perfectly clear. Trouble starts when the same cases appear across several operations. Adding one new type then means editing pricing, validation, formatting, scheduling, and other branches in separate places. The type list has become a change axis, but the code still represents it as scattered conditionals. Replacing type branches with polymorphism moves behavior for each variant behind a shared contract. Callers ask for an operation without selecting the implementation themselves. This can concentrate related rules and reduce repeated branching, but it also introduces more types and indirection. The refactoring pays off only when that trade is useful.

Software Engineering 11 Sep 2026 8 min read

Replacing Inheritance with Delegation

Replacing Inheritance with Delegation A class needs three useful methods from another class, so extending that class seems convenient. Months later, the subclass also inherits methods it shouldn’t expose, depends on initialization details it doesn’t control, and breaks when the superclass changes an internal assumption. The problem isn’t inheritance itself. The problem is using an is-a relationship to obtain code reuse when the real relationship is uses-a. Replacing inheritance with delegation makes that relationship explicit: the object keeps a collaborator and forwards only the behavior it actually needs.

Software Engineering 11 Sep 2026 8 min read

Replace Nested Conditionals with Guard Clauses

Replace Nested Conditionals with Guard Clauses A method often starts simple and becomes deeply nested one condition at a time. A null check wraps a permission check. That wraps a state check. The actual operation ends up several indentation levels from the method boundary, even though it is the main path a reader cares about. Guard clauses handle exceptional or disqualifying conditions near the top of a method and exit immediately. The remaining code can then describe the normal path with less structural noise.