MiniMax explains why its LLM failed to generate the name Ma Jiaqi
AIMiniMax says its M2 series could not output the name Ma Jiaqi, a failure it traced to post-training data that rarely included the token. Its tests found the input embedding stayed stable while the lm_head weights for low-frequency tokens drifted during SFT. A synthetic full-vocabulary repetition dataset restored generation for affected tokens and reduced Japanese-to-Russian confusion from 47% to 1%.
Why it matters: The post traces a community-noticed token failure through tokenizer, embedding, and lm_head tests, showing how post-training data coverage can cause low-frequency token drift.






