Skip to content

Archive

Load Shedding

3 articles
Software Engineering 20 Sep 2026 7 min read

Load Shedding Protects Useful Work When Capacity Is Exhausted

A service can be healthy at 2,000 requests per second and collapse at 2,400. The extra 400 requests do not merely wait their turn. They may occupy connection slots, queue entries, memory, worker threads, database sessions, and retry budgets while useful throughput falls. Load shedding places an explicit admission decision before a scarce resource is fully consumed. When the system cannot serve all incoming work within its operating envelope, it rejects selected requests early instead of allowing every request to compete until they all become slow.

Software Engineering 20 Sep 2026 7 min read

Bounded Queues Turn Overload into Explicit Rejection

Bounded Queues Turn Overload into Explicit Rejection A queue can absorb short bursts when requests arrive faster than workers can finish them. That buffer is useful only while it remains a buffer. If producers can keep adding work without a fixed limit, sustained overload turns the queue into an expanding inventory of requests that may wait long after their results are useful. A bounded queue changes the failure mode. It accepts waiting work up to a deliberate capacity, then refuses additional admission until space becomes available. The service still experiences overload, but the overload appears as an explicit control decision rather than unbounded growth in memory and waiting time.

Software Engineering 12 Sep 2026 11 min read

Load Shedding: Reject Work Before Overload Spreads

Load Shedding: Reject Work Before Overload Spreads A service has finite capacity. When offered work exceeds that capacity for long enough, accepting every request can make the service less useful rather than more useful. Queues grow, deadlines expire, memory pressure rises, dependencies receive more traffic, and successful throughput can fall. Load shedding is the deliberate rejection of work that the system cannot serve within an acceptable budget. The goal is not to maximize the number of requests admitted. The goal is to preserve useful service during overload.