A sparsely activated Mixture-of-Experts layer can keep many expert parameters while sending each token to only a small subset of them. Once those experts are partitioned across devices, sparse activation does not imply local execution. A token assigned to a remote expert has to move to the device that owns that expert, and its expert output has to return to the device that continues the model computation.

This movement is the defining systems cost of expert parallelism. The router makes a logical selection, but a distributed runtime must turn that selection into communication, local expert batches, expert computation, and a reverse exchange.

Routing creates a data-placement problem

Consider a layer with tokens initially distributed across several devices and experts sharded across the same group. The router computes expert assignments for local tokens. Some assignments target local experts; others target experts owned elsewhere.

The runtime therefore groups token representations by destination. Conceptually, the forward path has this shape:

local token states
      |
router assignments
      |
pack by expert owner
      |
all-to-all dispatch
      |
local expert batches
      |
expert computation
      |
all-to-all return
      |
restore token order

The collective is not part of the mathematical definition of every MoE architecture. It follows from a particular placement strategy: experts are distributed across devices while tokens can select experts outside their current device. A deployment that replicates every expert locally has a different communication pattern and a much larger parameter-memory cost.

Dispatch volume follows routed token state, not total expert parameters

Expert parallelism separates parameter placement from activation traffic. A device does not send an expert’s full weight matrix for each token. Instead, it sends token state needed by the remote expert and receives the resulting state back.

For a fixed hidden-state width, communication volume is therefore tied to the number of routed token copies, their representation size, routing multiplicity, and the participating device layout. Top-2 routing can create more token destinations than top-1 routing because one token may be dispatched to two experts.

This distinction matters when comparing sparse compute with communication. Activating fewer experts can reduce arithmetic relative to evaluating every expert, yet remote activation can still produce substantial network traffic. Parameter sparsity and communication sparsity are related only through the concrete routing and placement scheme.

Expert imbalance becomes communication imbalance

Routers do not necessarily assign equal token counts to every expert. If many tokens select experts owned by one device, that device receives a larger expert batch while other devices may receive less work.

The effect is visible in both computation and communication. A heavily selected destination must receive more token state, process more expert work, and return more outputs. Collective completion can then be constrained by the slowest or most heavily loaded participant rather than the average participant.

Capacity limits, auxiliary balancing objectives, routing policies, and token-dropping or rerouting rules can alter this behavior. Their exact semantics are architecture-specific. A balanced expert count in the model definition does not itself guarantee balanced traffic at runtime.

All-to-all exposes the physical interconnect

An all-to-all exchange describes a communication relationship, not a uniform hardware cost. Transfers within one accelerator package, across a high-bandwidth node fabric, and across nodes through a network can have very different latency and bandwidth characteristics.

As a result, the same expert-parallel degree can behave differently under different placement topologies. Mapping frequently communicating expert groups inside a faster domain can reduce expensive cross-domain traffic. Extending an expert-parallel group across slower links can expose communication that was less significant on a single node.

This is also a boundary on portable performance claims. A throughput result from one accelerator count, topology, collective library, message size, and routing distribution does not establish the same ratio on another system.

Packing and permutation are part of the execution path

Before dispatch, routed tokens usually have to be rearranged into buffers suitable for communication and expert execution. After expert outputs return, the runtime has to associate them with the original token positions and combine multiple expert contributions when the routing rule requires it.

Those transformations consume memory bandwidth and may require temporary buffers. They can also affect kernel shapes: balanced routing can produce reasonably sized expert batches, while fragmented routing can leave many small batches that use compute resources differently.

An implementation can fuse, overlap, or reorganize parts of this path, but those are runtime properties rather than guarantees of the MoE abstraction. The relevant execution cost includes routing metadata, permutation, collective communication, expert kernels, reverse communication, and restoration of token order.

Communication can overlap with compute only when dependencies permit it

Distributed runtimes may attempt to overlap token exchange with expert computation. For example, chunks that have already arrived can sometimes begin processing while other chunks are still in flight. The achievable overlap depends on scheduling, buffer ownership, collective implementation, expert batch shape, and hardware concurrency.

Overlap does not remove transferred bytes. It changes how much communication latency remains exposed on the critical path. If expert computation is short relative to transfer time, or if communication and computation contend for the same resources, a large fraction of the exchange can remain visible.

For that reason, statements that expert parallelism is compute-efficient and statements that it is communication-efficient are separate claims. Sparse expert activation addresses arithmetic selection. Distributed placement determines the movement required to realize that selection.

Placement defines the practical scaling boundary

Adding experts can increase total model capacity without activating all expert parameters for every token. Expert parallelism makes that parameter set fit across multiple devices, but the scaling path introduces a second resource dimension: routed token traffic.

The practical boundary is set by the joint system. Router distribution determines destinations, expert placement maps those destinations to devices, collectives move token states, and the interconnect determines the cost of that movement. Compute capacity alone does not characterize the layer.

Expert parallelism is therefore best treated as a distributed execution scheme for sparse routing, not merely a way to split expert weights. Its central trade is explicit: sharding experts reduces per-device expert-parameter residency, while remote expert selection turns routing decisions into communication that must be scheduled, balanced, and carried by the physical topology.