Skip to content
Read the original: Fireworks AI Blog· Published 51/100AI score51/100

Fireworks explains how numerical mismatch and MoE routing can derail RL training

Original titleReinforcement learning: Why alignment of numerics and MoE routing matter

AISummary

Numerical differences between a rollout engine and a trainer can make reinforcement learning collapse even when algorithm and data stay identical.

In a GLM 5.2 experiment, reward fell from about 0.9 to under 0.2 around step 20 without alignment, while aligned numerics kept reward stable over 25 steps.

A Qwen3.5-MoE investigation traced a significant mismatch to how expert outputs were combined, and router replay alone was judged insufficient.

Read the original fireworks.ai

Source: Fireworks AI Blog · fireworks.aiPublished · added here