Skip to content

Companies & models · Latest news

Kimi / Moonshot AI

Follow Kimi models and products from Moonshot AI, including open K-series models, long-context work, and product changes.

9 top picks · 1 in the past 30 days · chosen from 38 items collected

Latest pick Key moments

Kimi / Moonshot AI top picks

Sep 22

Sep 22Tue
  1. Fireworks AI BlogAI score65

    Fireworks releases Ember-1, a Kimi K3 variant that cuts reasoning tokens by about 40%

    AIFireworks Research released Ember-1, a specialized model built on Kimi K3 that it says delivers the same quality with 40% fewer tokens. Across five industry benchmarks, Ember-1 matched K3 max quality at a fraction of the cost, and in two customer A/B tests it used about 35% fewer tokens per task. It is available as a Research Preview on Serverless, and Fireworks is also launching training support for customized models.

    Why it matters: The source gives benchmark and A/B results for cutting reasoning tokens while holding quality, which bears on cost planning for coding and agent workloads.

Jul 27

Jul 27Mon
  1. KimiAI score86

    Moonshot AI releases Kimi K3 weights and technical report

    AIMoonshot AI is releasing the model weights and technical report for Kimi K3, a 2.8T-parameter MoE model with native visual understanding and a 1M-token context window. The post says the new architecture delivers 2.5x the intelligence per unit of compute, and the company is also opening high-performance attention kernels, an MoE communication library, and infrastructure for running agent environments at scale.

    Why it matters: The source names the model size, context window, and released weights, which helps readers compare its scale and openness with other frontier releases.

Jun 13

Jun 13Sat
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score88

    Moonshot AI releases open-weight Kimi K3 with 2.8T parameters and 1M context

    AIMoonshot AI released Kimi K3 on Hugging Face as an open-weight, native multimodal agentic model with 2.8T total parameters and 104B activated parameters. It supports a 1-million-token context window and text and image input, with weights released under the Kimi K3 License. The model card reports benchmark results for coding, agentic, and vision tasks against several closed models, and recommends vLLM, SGLang, or TokenSpeed for inference.

    Why it matters: The release pairs open weights with a 2.8T-parameter MoE architecture and benchmark tables against several named closed models, useful for comparing frontier capability claims.

Jun 11

Jun 11Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score62

    Moonshot AI releases Kimi K2.7 Code, a coding-focused agentic model

    AIMoonshot AI published Kimi-K2.7-Code, a coding-focused agentic model built on Kimi K2.6, with a 1T-parameter MoE architecture and 32B activated parameters. The model card reports about 30% fewer thinking tokens than K2.6 and benchmark results against GPT-5.5 and Claude Opus 4.8, with weights and code released under a Modified MIT License.

    Why it matters: The model card gives benchmark comparisons against GPT-5.5 and Claude Opus 4.8 on coding and agentic tasks, useful for judging its position among current coding models.

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Jan 1

Jan 1Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score75

    Moonshot AI releases open-source multimodal agent model Kimi K2.5

    AIMoonshot AI released Kimi K2.5, an open-source native multimodal agentic model built by continual pretraining on about 15 trillion mixed visual and text tokens. The model card reports a 1T-parameter Mixture-of-Experts architecture with 32B activated parameters and a 256K context length, and it lists benchmark results against GPT-5.2, Claude 4.5 Opus, Gemini 3 Pro, DeepSeek V3.2, and Qwen3-VL-235B-A22B-Thinking. Weights and code are released under a Modified MIT License, with API access on the Moonshot platform.

    Why it matters: The model card gives a full benchmark table against GPT-5.2, Claude 4.5 Opus, and Gemini 3 Pro, useful for comparing open multimodal agent models.

Nov 4, 2025

Nov 4, 2025Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score82

    Moonshot AI releases open-source Kimi K2 Thinking reasoning agent model

    AIMoonshot AI released Kimi K2 Thinking, an open-source thinking model that interleaves step-by-step reasoning with tool calls across 200 to 300 sequential invocations. The model is a 1T-parameter mixture-of-experts with 32B activated parameters and a 256k context window, and it uses native INT4 quantization for roughly 2x faster generation. The model card reports benchmark results on HLE, BrowseComp, and other tests, and recommends vLLM, SGLang, or KTransformers for deployment.

    Why it matters: The model card gives benchmark tables, quantization details, and deployment settings, letting readers compare Kimi K2 Thinking against GPT-5 and other models on specific tasks.

Oct 30, 2025

Oct 30, 2025Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score60

    Moonshot AI releases Kimi Linear 48B hybrid linear attention models on Hugging Face

    AIMoonshot AI released Kimi Linear, a hybrid linear attention architecture with 48B total and 3B activated parameters and a 1M-token context length, on Hugging Face. The model card reports up to 6.3x faster TPOT than MLA at 1M tokens and up to 75% lower KV cache needs, and says it outperforms full attention on long-context and RL-style benchmarks.

    Why it matters: The model card gives concrete long-context speed and memory figures for a hybrid attention design, useful for judging whether linear attention can replace full attention in practice.

  2. Moonshot AI (Kimi) · new models on Hugging FaceAI score72

    Moonshot AI releases Kimi Linear 48B-A3B hybrid attention models on Hugging Face

    AIMoonshot AI has released Kimi-Linear-Base and Kimi-Linear-Instruct, both 48B total and 3B activated parameters with a 1M context length, on Hugging Face. The models use Kimi Delta Attention in a 3:1 hybrid ratio with global MLA, cutting KV cache by up to 75% and boosting decoding throughput by up to 6x at 1M tokens. The KDA kernel is open-sourced in FLA, and the checkpoints were trained on 5.7T tokens.

    Why it matters: The model card gives concrete throughput and KV cache figures for a hybrid attention design, which helps readers weigh its long-context tradeoffs against full attention.

Key moments

Since 2023
  1. ProductKimi chatbot launches with a long context window
  2. ProductKimi extends context to 2 million Chinese characters
  3. ModelKimi k1.5 reasoning model released
  4. ModelKimi K2 released as a trillion-parameter open model
  5. ModelKimi K2 Thinking released
  6. ModelMoonshot AI releases open-source multimodal agent model Kimi K2.5
  7. ModelMoonshot AI releases open-source Kimi K2.6 multimodal agentic model
  8. ModelMoonshot AI releases open-weight Kimi K3 with 2.8T parameters and 1M context
  9. ModelKimi K3 released with open weights