Using Prime Intellect, Extropic post-trained Qwen3.6-35B-A3B for thermodynamic ML research, nearly tripling its eval results on held-out tasks in ~100 GRPO steps.
They built a custom RL environment with verifiers and trained on Hosted Training, Prime Sandboxes, and Prime Inference. This meant Extropic didn't have to manage multi-node GPU infrastructure, so the team could focus on research.

