MiniMax releases M3-MXFP8, a 1M-context native multimodal model on Hugging Face
Original titleMiniMaxAI/MiniMax-M3-MXFP8
AISummary
MiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters.
M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context.
The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.
AIWhy it matters
The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.
Source: MiniMax · new models on Hugging Face · huggingface.coPublished · added here