Qwen3.6 reward rises 2.8x via GRPO on Hosted Training
Original titleAfter ~100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain.
AISummary
Prime Intellect reports that after about 100 GRPO steps on Hosted Training, Qwen3.6's reward on held-out problems rose from 0.127 to 0.361, a 2.8x gain. Qwen3.5, trained the same way, reached 0.356, suggesting the method works across model families.
Both post-trained models finished well ahead of other open models and narrowed the gap to Claude Opus 4.8, with Qwen3.6 activating only 3B parameters per token.
Source: Prime Intellect · x.comPublished · added here