Prime Intellect rebuilds GLM-5.2 RL weight transfer on NIXL, cutting sync to 3.9 seconds
Original titleGLM-5.2 RL weight transfer in 4 seconds using NIXL and ModelExpress
AISummary
Prime Intellect reports that rebuilding RL weight transfer for GLM-5.2 on NIXL and ModelExpress cut sync time from 86.1 seconds with NCCL to 3.9 seconds in its fastest setting.
The method traces vLLM's loader to find each tensor's runtime layout, then reads only the needed source bytes over RDMA and replays the rest locally.
Most remaining latency comes from vLLM's pause consensus, which the team reduced by syncing every wave instead of every 32.
Source: Prime Intellect Blog · primeintellect.ai