Token shuffling routes tokens to predicted experts without extra network traffic
AIZyphra reports that expert routing in mixture-of-experts models is predictable across layers, since the experts a token uses in one layer indicate which it will need next. Its token shuffling method moves each token to the GPU holding those experts within a transfer that already runs after attention, adding no network traffic.