Zyphra reports up to 2.63x faster MoE token exchange in Megatron-LM
Original titleThe gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes...
AISummary
Zyphra reports that its MoE training optimizations speed up token exchange by 1.16x to 2.63x and full training steps by up to 1.41x in Megatron-LM on 8 to 64 GPUs. The gains are largest when each token uses more experts and those experts span several nodes.
Source: Zyphra · x.comPublished · added here