Zyphra's correlated expert placement cuts MoE token copies by 58%
AIZyphra places experts that are frequently chosen together on the same GPU, so each token is sent to a GPU only once regardless of how many of its experts reside there. On 8 GPUs, this removes up to 58% of token copies sent across the network.








