Skip to content
View original post on X: Prime Intellect· 32/100AI score32/100

Extropic uses Prime Intellect to post-train Qwen3.6-35B-A3B for thermodynamic ML

AISummary

Extropic post-trained Qwen3.6-35B-A3B with Prime Intellect for thermodynamic ML research, nearly tripling its held-out eval results in about 100 GRPO steps. The team built a custom RL environment with verifiers and trained on Hosted Training, Prime Sandboxes, and Prime Inference. This let Extropic avoid managing multi-node GPU infrastructure and focus on research.

Post on XView on X
@PrimeIntellect

Using Prime Intellect, Extropic post-trained Qwen3.6-35B-A3B for thermodynamic ML research, nearly tripling its eval results on held-out tasks in ~100 GRPO steps.

They built a custom RL environment with verifiers and trained on Hosted Training, Prime Sandboxes, and Prime Inference. This meant Extropic didn't have to manage multi-node GPU infrastructure, so the team could focus on research.

Extropic@extropic
What evolves together fits together. We’re using RL to post-train models for Thermo AI research, accelerating algorithmic discovery for our new species of computer. The first sparks of Thermo RSI, in collaboration with @PrimeIntellect. https://extropic.ai/writing/baby-thermo-rsi
View quoted post on X

Source: Prime Intellect · x.comPublished · added here