Read the original: OpenBMB (MiniCPM) · new models on Hugging Face· Published · added 45/100AI score45/100
openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning
Original titleopenbmb/JustRL-II-base-model
AISummary
OpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning.
The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline.
The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.
Source: OpenBMB (MiniCPM) · new models on Hugging Face · huggingface.coPublished · added here