Skip to content
View original post on X: Reflection· 23/100AI score23/100

Reflection AI's Beam model pretrained in four weeks on 24T tokens

AISummary

Reflection AI says its Beam model was pretrained in 4 weeks on 24T high-quality tokens, giving it innate coding capabilities.

The company credits MoE stability improvements and large-scale data curation and deduplication for a base model it claims outperforms open-source base models of the same class.

It presents this strong reasoning foundation as what makes sustained reinforcement learning gains possible.

Post on XView on X
@reflection_ai

A reply · the post it answers

Sustained RL improvement was made possible by a strong foundation for reasoning.

Beam was pretrained in 4 weeks on 24T high-quality tokens to have innate coding capabilities.

Innovations in MoE stability and large-scale data curation & deduplication enabled us to produce a foundation that outperforms available open-source base models of the same class.

Source: Reflection · x.comPublished · added here