How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72
Original titleHow we trained the fastest DSpark for Kimi-K3 using GB300 NVL72
The vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware.
They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines.
The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.
The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.
Source: vLLM Blog · vllm.aiPublished · added here