Zyphra speeds MoE token routing up to 2.63x on AMD MI300X GPUs
Original titleIn Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime.
AISummary
Zyphra researchers report that exploiting patterns in how tokens are routed to experts makes Mixture of Experts communication up to 2.63x faster on AMD MI300X GPUs. The speedup comes without changing the model itself.
Source: Zyphra · x.comPublished · added here