Moonshot AI releases Kimi Linear 48B-A3B hybrid attention models on Hugging Face
Original titlemoonshotai/Kimi-Linear-48B-A3B-Instruct
AISummary
Moonshot AI has released Kimi-Linear-Base and Kimi-Linear-Instruct, both 48B total and 3B activated parameters with a 1M context length, on Hugging Face.
The models use Kimi Delta Attention in a 3:1 hybrid ratio with global MLA, cutting KV cache by up to 75% and boosting decoding throughput by up to 6x at 1M tokens. The KDA kernel is open-sourced in FLA, and the checkpoints were trained on 5.7T tokens.
AIWhy it matters
The model card gives concrete throughput and KV cache figures for a hybrid attention design, which helps readers weigh its long-context tradeoffs against full attention.
Source: Moonshot AI (Kimi) · new models on Hugging Face · huggingface.co