Expert Parallelism Turns MoE Routing into All-to-All Communication
A sparsely activated Mixture-of-Experts layer can keep many expert parameters while sending each token to only a small subset of them. Once those experts are partitioned across devices, sparse activation does not imply local execution. A token assigned to a remote expert has to move to the device that owns that expert, and its expert output has to return to the device that continues the model computation. This movement is the defining systems cost of expert parallelism. The router makes a logical selection, but a distributed runtime must turn that selection into communication, local expert batches, expert computation, and a reverse exchange.