Skip to content

Areas · Latest news

Open-source ecosystem

Open models, frameworks, and repositories: open weights, breakout community projects, and the balance between open and closed AI.

145 top picks · 63 in the past 30 days · chosen from 1,155 items collected

Latest pick

Top picks archive · Page 8

Top picks 141–145 of 145

Nov 4, 2025

Nov 4, 2025Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score82

    Moonshot AI releases open-source Kimi K2 Thinking reasoning agent model

    AIMoonshot AI released Kimi K2 Thinking, an open-source thinking model that interleaves step-by-step reasoning with tool calls across 200 to 300 sequential invocations. The model is a 1T-parameter mixture-of-experts with 32B activated parameters and a 256k context window, and it uses native INT4 quantization for roughly 2x faster generation. The model card reports benchmark results on HLE, BrowseComp, and other tests, and recommends vLLM, SGLang, or KTransformers for deployment.

    Why it matters: The model card gives benchmark tables, quantization details, and deployment settings, letting readers compare Kimi K2 Thinking against GPT-5 and other models on specific tasks.

Oct 30, 2025

Oct 30, 2025Thu
  1. Moonshot AI (Kimi) · new models on Hugging FaceAI score60

    Moonshot AI releases Kimi Linear 48B hybrid linear attention models on Hugging Face

    AIMoonshot AI released Kimi Linear, a hybrid linear attention architecture with 48B total and 3B activated parameters and a 1M-token context length, on Hugging Face. The model card reports up to 6.3x faster TPOT than MLA at 1M tokens and up to 75% lower KV cache needs, and says it outperforms full attention on long-context and RL-style benchmarks.

    Why it matters: The model card gives concrete long-context speed and memory figures for a hybrid attention design, useful for judging whether linear attention can replace full attention in practice.

  2. Moonshot AI (Kimi) · new models on Hugging FaceAI score72

    Moonshot AI releases Kimi Linear 48B-A3B hybrid attention models on Hugging Face

    AIMoonshot AI has released Kimi-Linear-Base and Kimi-Linear-Instruct, both 48B total and 3B activated parameters with a 1M context length, on Hugging Face. The models use Kimi Delta Attention in a 3:1 hybrid ratio with global MLA, cutting KV cache by up to 75% and boosting decoding throughput by up to 6x at 1M tokens. The KDA kernel is open-sourced in FLA, and the checkpoints were trained on 5.7T tokens.

    Why it matters: The model card gives concrete throughput and KV cache figures for a hybrid attention design, which helps readers weigh its long-context tradeoffs against full attention.

May 21, 2025

May 21, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition launches official DeepWiki MCP server for indexed GitHub repos

    AICognition launched the official DeepWiki Model Context Protocol server, which is free and requires no login or authentication. It gives programmatic access to ask_question, read_wiki_contents, and read_wiki_structure for GitHub repositories indexed on DeepWiki.com. Private repositories require a Devin account with GitHub connected, and open-source maintainers can apply for $500 in Devin credits.

    Why it matters: The source names the three tools and the access path, showing how indexed GitHub repositories can be queried programmatically without login.

Nov 30, 2024

Nov 30, 2024Sat
  1. Liquid AI BlogAI score60

    Liquid AI's STAR uses evolutionary search to synthesize tailored model architectures

    AILiquid AI reports STAR, an evolutionary algorithm that synthesizes tailored neural network architectures from a numerical genome representation. The authors say it produced hundreds of designs that outperform Transformer and hybrid architectures in quality, with smaller caches and parameter counts, and can optimize for latency on target hardware. The full method is described in the arXiv technical report 2411.17800.

    Why it matters: The post explains how evolutionary search over a new architecture design space produced designs beating Transformers and hybrids, giving a concrete method for quality versus latency and memory trade-offs.