Rigel: 2.3B hybrid Mamba-2 MoE nears Llama-3.2-3B with <1% FLOPs
Original titleMayank put in a crazy amount of work to get pretraining to work on 3 gens of Nvidia GPUs and 2 gens of TPUs! Very good model for such a s...
AISummary
Mayank's Rigel, a 2.3B-parameter MoE (360M active) hybrid Mamba-2 model, was pretrained across H100, A100, V100 GPUs and TPU v5p/v6e on one codebase. The model lands within a few points of Llama-3.2-3B while using under 1% of its pretraining FLOPs. Tri Dao praised the work's engineering effort and the model's strength for its small size.
Source: Tri Dao · x.comPublished · added here