Token shuffling routes tokens to predicted experts without extra network traffic
Original titleRouting is also predictable across layers: the experts a token used in one layer tell us which it will likely need next. Token shuffling ...
AISummary
Zyphra reports that expert routing in mixture-of-experts models is predictable across layers, since the experts a token uses in one layer indicate which it will need next. Its token shuffling method moves each token to the GPU holding those experts within a transfer that already runs after attention, adding no network traffic.
Source: Zyphra · x.comPublished · added here