Switch Routing Capacity Turns Expert Imbalance into Token Overflow
A Switch-style sparse layer can send many tokens toward the same expert even when every expert has identical compute capacity. The router makes token-level choices from model-produced scores; it does not inherently produce an even partition. A fixed expert capacity therefore creates a hard boundary between routing preference and the amount of expert computation admitted for a batch. In the Switch Transformer formulation, each token is routed to the expert with the highest router probability. Each expert receives a fixed token capacity derived from the token count, expert count, and a capacity factor. When assignments exceed that capacity, excess tokens overflow instead of enlarging the expert batch without bound.