vLLM Reports PD Serving Results for Qwen3.8-2.4T on GB300 NVL72
AIvLLM achieved 5000 total token throughput per GPU in high-throughput PD serving of Qwen3.8-2.4T on a GB300 NVL72 cluster under an 8K/1K workload. The low-latency scenario reached 180 generated tokens per user, with both results shown on the Pareto frontier. The post also provides srt-slurm recipes and explains the tuning process used to create them.




