Ai2 releases Olmo-core 3, an open framework for training large MoE models
Original titleIntroducing Olmo-core 3: Open, scalable training infrastructure for large MoEs
AISummary
Ai2 released Olmo-core 3, an open training framework redesigned to scale mixture-of-experts models into the trillion-parameter range.
In one benchmark, expert count rose from 8 to 128 with about 3.2B active parameters per token, total capacity grew from 4.6B to 47B, and throughput fell by less than 5%.
The framework is fully open, so researchers can train their own MoEs and experiment with routing and parallelism.
AIWhy it matters
The release documents concrete MoE scaling results and reported failure modes, useful for teams weighing training-stack tradeoffs before adopting an open framework.
Source: Ai2 (Allen Institute for AI) · allenai.org