Skip to content

Archive

Scheduling

5 articles
Artificial Intelligence 24 Sep 2026 5 min read

Iteration-Level Scheduling Rebuilds LLM Batches Between Decode Steps

Autoregressive serving does not give every request the same completion point. One sequence may emit an end token after a few decode steps while another remains active for hundreds. If the serving engine keeps the original request batch fixed until every member finishes, completed sequences leave execution capacity stranded behind longer sequences. Iteration-level scheduling moves the scheduling boundary inward. Instead of treating an entire request as the indivisible scheduling unit, the engine returns to the scheduler after a model iteration. Finished sequences can leave, waiting work can enter, and the next iteration can run with a different set of active sequences.

Artificial Intelligence 24 Sep 2026 5 min read

Continuous Batching Replaces Static Request Groups with Iteration-Level Scheduling

Autoregressive generation does not finish every request at the same iteration. One sequence may emit a stop token after a few decode steps while another remains active for hundreds more. A static batch keeps those requests coupled until the batch boundary. Continuous batching breaks that coupling by letting the active set change between model iterations. The key change is scheduling granularity. A request is no longer the indivisible scheduling unit for the entire generation. The serving system can construct an execution batch for one iteration, update request state after that iteration, remove finished sequences, and admit queued work before the next execution step.

Software Engineering 23 Sep 2026 5 min read

Priority Inheritance Bounds Priority Inversion Around Mutexes

Priority Inheritance Bounds Priority Inversion Around Mutexes Priority scheduling does not guarantee that the highest-priority runnable task can always make progress. A high-priority task can block on a mutex held by a lower-priority task. If medium-priority work then preempts the owner, the high-priority task remains blocked even though the medium-priority work has no direct dependency on the mutex. This is priority inversion. The inversion starts with an ordinary dependency: the high-priority task needs a resource owned by a lower-priority task. The damaging part is interference from tasks between those priorities, which can delay the owner and extend the blocking interval.

Software Engineering 20 Sep 2026 5 min read

Power of Two Choices Reduces Load Imbalance with Two Samples

Power of Two Choices Reduces Load Imbalance with Two Samples A load balancer that chooses one destination uniformly at random is cheap and decentralized, but random placement can produce uneven queues. At the other extreme, selecting the least loaded destination from the entire pool requires current load information for every candidate and can make the balancer itself expensive. The power-of-two-choices strategy sits between those designs. For each request, sample two eligible destinations, compare a load signal, and send the request to the better candidate. Two observations are enough to avoid many unlucky placements without requiring a global search.

Artificial Intelligence 19 Sep 2026 8 min read

Chunk Long Prefills to Limit Decode Stalls in LLM Serving

A long prompt can occupy an accelerator for a much larger scheduling interval than a single decode iteration. When a serving engine mixes new prefills with requests that are already generating tokens, that difference can show up as irregular time between output tokens. The model has not changed; the interference comes from how two distinct inference phases share execution time. Prefill processes a prompt and builds the key-value state required by later causal attention. Decode then extends the sequence autoregressively, usually one new token per active request per iteration. Those phases place different pressure on hardware, so treating them as interchangeable scheduling units can produce avoidable stalls.