MiniMax releases M3, a native multimodal model with 1M context
Original titleMiniMaxAI/MiniMax-M3
AISummary
MiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters.
The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context.
M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.
AIWhy it matters
The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.
Source: MiniMax · new models on Hugging Face · huggingface.coPublished · added here