Skip to content
Read the original: vLLM Blog· Published Pick62/100AI score62/100

How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

Original titleHow we trained the fastest DSpark for Kimi-K3 using GB300 NVL72

AISummary

The vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware.

They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines.

The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

AIWhy it matters

The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

Read the original vllm.ai

Source: vLLM Blog · vllm.aiPublished · added here