Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

Jun 4

Jun 4Thu
  1. Georgi GerganovXAI score44

    llama.cpp adds multi-GPU and tensor parallel support with NVIDIA

    AIMaintainers and NVIDIA engineers improved multi-GPU performance in ggml, the low-level engine behind llama.cpp, yielding significant gains on RTX systems. The work also lays groundwork for hardware-agnostic tensor parallelism in ggml. Details are in a technical blog from NVIDIA RTX Spark.

    Image from @ggerganov's post

Jun 3

Jun 3Wed
  1. PaddlePaddleOfficialAI score31

    Baidu CoBuddy, a free code-focused model, now live on Novita

    AIBaidu CoBuddy is now available for free on Novita AI, a code-focused model aimed at developers and AI agents. It offers a 131K context window, up to 65K output tokens, and native tool calling. The model is served through Novita's serverless API for high-throughput, low-latency inference.

Jun 2

Jun 2Tue
  1. MiniMax · new models on Hugging FaceOfficialAI score78

    MiniMax releases M3-MXFP8, a 1M-context native multimodal model on Hugging Face

    AIMiniMax published MiniMax-M3-MXFP8, an MXFP8 quantized variant of its native multimodal M3 model with 1M context, about 428B total parameters and about 23B activated parameters. M3 adds MiniMax Sparse Attention, which the source says yields 9× prefill and 15× decode speedups over M2 at 1M context. The model supports three thinking modes (enabled, adaptive, disabled) via the thinking parameter and can be served with SGLang, vLLM, or Transformers.

    Why it matters: The release pairs sparse attention for 1M-token contexts with reported prefill and decode speedups over M2, useful for judging long-context serving costs.

  2. MiniMax · new models on Hugging FaceOfficialAI score68

    MiniMax releases M3, a native multimodal model with 1M context

    AIMiniMax has released MiniMax-M3, a native multimodal model with a 1M-token context window, roughly 428B total parameters, and about 23B activated parameters. The model introduces MiniMax Sparse Attention, which the source says delivers 9× prefill and 15× decode speedups over M2 at 1M context. M3 supports enabled, adaptive, and disabled reasoning modes through the thinking parameter, and weights are available on Hugging Face.

    Why it matters: The source gives concrete attention-efficiency figures and three reasoning modes, which helps readers judge long-context cost against deployment choices.

  3. ByteDance · new models on Hugging FaceOfficialAI score44

    ByteDance Releases Bernini-R Diffusers Weights for Video Generation and Editing

    AIByteDance has open-sourced the inference code and model weights of the Bernini Renderer (Bernini-R), a DiT-based renderer paired with an MLLM-based semantic planner for video generation and editing. A diffusers-format version, ByteDance/Bernini-R-Diffusers, bundles the Wan2.2 base components with the Bernini-R transformer weights for direct loading, and the framework requires a CUDA GPU with PyTorch 2.5.1+cu124.

Jun 1

Jun 1Mon
  1. Cognition Blog (Devin, Windsurf)OfficialAI score50

    Cognition launches Devin Desktop, the next generation of Windsurf

    AICognition has announced Devin Desktop, the next generation of Windsurf, which makes the Agent Command Center the default IDE surface for managing local and cloud agents, PRs, and context. Spaces let related agents share context, and Agent Client Protocol (ACP) support lets any ACP-compatible agent run alongside Devin. The IDE remains fully backwards-compatible with Windsurf, including editor extensions, keybindings, LSPs, and terminal workflows.

  2. PaddlePaddleOfficialAI score36

    PaddleOCR and ERNIE Image now available as official Dify plugins

    AIPaddleOCR and ERNIE Image are now available as official Dify plugins, bringing document parsing and image generation into Dify's agent workflows. PaddleOCR, powered by PP-OCRv5, PP-StructureV3, and PaddleOCR-VL, turns images, scanned PDFs, and multilingual documents into structured data for chunking, vectorization, and RAG, with private or on-prem deployment supported. ERNIE Image offers free generation, a Turbo mode with 8-step inference, and an OpenAI-style API.

    Image from @PaddlePaddle's post

May 31

May 31Sun
  1. MiniMax BlogOfficialAI score82

    MiniMax M3 releases with 1M context, native multimodality and sparse attention

    AIMiniMax released M3, an open-weight model with a 1M-token context window, native image and video input, and desktop operation support. The post credits a new sparse attention architecture, MSA, for long-context gains, reporting over 9x prefilling and over 15x decoding speedups and 59.0% on SWE-Bench Pro. The API and MiniMax Code are available now, with the technical report and open weights promised within 10 days.

    Why it matters: The post pairs a new sparse attention design with benchmark figures and a 1M-token context window, letting readers judge the architecture's practical effect on long-context work.

May 30

May 30Sat
  1. Xiaomi MiMoOfficialAI score62

    Xiaomi details how it turned MiMo-V2.5 Hybrid SWA savings into production inference gains

    AIXiaomi describes an end-to-end inference optimization for the MiMo-V2.5 series, centered on Hybrid SWA, which it says cuts KVCache storage to roughly 1/7 of Full Attention. The post covers a dual KVCache pool design, SWA-aware prefix cache matching, the GCache distributed cache, and scheduling changes, and reports cache hit rates averaging 93% in server-side observations. It also covers prefill and decode optimizations, multimodal encoder improvements, and open-source contributions to SGLang.

    Why it matters: The post explains how Hybrid SWA's theoretical KVCache savings were realized in production through dual pools, SWA-aware prefix caching, and tiered storage, giving concrete engineering patterns for long-context inference.

May 28

May 28Thu
  1. PaddlePaddleOfficialAI score36

    PaddleOCR-VL 1.6 released with 96.33% SOTA on OmniDocBench

    AIPaddlePaddle has released PaddleOCR-VL 1.6, which sets a new state-of-the-art score of 96.33% on OmniDocBench for text, formula, and table recognition. It ranks first on OmniDocBench v1.5 and Real5-OmniDocBench, with gains in table, classic text, rare character, seal, spotting, and chart recognition. The version is fully compatible with the v1.5 architecture, requiring no migration.

    Image from @PaddlePaddle's post

May 26

May 26Tue

May 24

May 24Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceOfficialAI score45

    Fun-ASR-Nano-2512-hf: Alibaba's Speech Recognition Model Gets Transformers Version

    AIFunAudioLLM has released Fun-ASR-Nano-2512-hf, a Hugging Face Transformers-compatible version of its end-to-end speech recognition model, which supports Chinese, English, and Japanese. The Chinese coverage includes 7 dialect groups and 26 regional accents, and a separate Fun-ASR-MLT-Nano-2512 checkpoint handles 31-language recognition. Developers can run the model natively in Transformers 5.17.0 without custom model code or trust_remote_code=True.

May 22

May 22Fri
  1. ReflectionOfficialAI score38

    Reflection signs MOU with US Department of Energy for national labs

    AIReflection, an open-source AI lab, has signed a memorandum of understanding with the US Department of Energy to explore strategic collaborations under the Genesis Mission. Through the partnership, Reflection will provide open-weight models to the DOE's 17 National Laboratories, customizable on lab-specific scientific data and deployed into researcher workflows. The Genesis Mission focuses on advanced nuclear, fusion and grid modernization, quantum ecosystem growth, and national security AI.

May 20

May 20Wed
  1. Stability AIOfficialAI score62

    Stability AI releases Stable Audio 3.0 model family with open-weight music models

    AIStability AI released Stable Audio 3.0, a family of four audio models trained on fully licensed data. Three of them, Small SFX, Small and Medium, have open weights on Hugging Face, while Large is available through the Stability AI API and enterprise self-hosting. Outputs can be distributed and commercialized under the Stability AI Community License, and organizations with more than $1M in annual revenue can use the Enterprise License.

    Why it matters: The source specifies which models are open-weight, their licensing terms, and clip-length limits, which matters for anyone deciding whether to build on them.

  2. PaddlePaddleOfficialAI score36

    PaddleOCR 3.5 adds Hugging Face Transformers as inference backend

    AIPaddleOCR 3.5 now supports Hugging Face Transformers as an inference backend, letting users run PP-OCRv5 and PaddleOCR-VL 1.5 models directly within the Transformers ecosystem. Users can select it with engine="transformers" while keeping the same PaddleOCR pipeline, which the post says eases integration for RAG and Document AI applications.

May 19

May 19Tue
  1. Awni HannunXAI score46

    PyTorch models can now run on Apple silicon via MLX delegate

    AIAwni Hannun highlights a workflow to write PyTorch models, export them, and run them with MLX on Apple silicon. The linked PyTorch post says ExecuTorch now includes an MLX delegate that runs PyTorch models on Apple silicon GPUs, supporting LLMs, speech-to-text, and MoE models.

    Image from @awnihannun's post
  2. koray kavukcuogluXAI score72

    Google rolls out Gemini 3.5 Flash globally across consumer, developer, and enterprise platforms

    AIGoogle is rolling out Gemini 3.5 Flash globally for consumers in the Gemini app and Search AI Mode. It is also available to developers through the Gemini API, Google Antigravity, and Google AI Studio, and to businesses on the Gemini Enterprise Agent Platform.

    Why it matters: The post shows where each Gemini 3.5 Flash access path goes, from consumer apps to developer and enterprise platforms, which helps readers pick the right entry point.

May 17

May 17Sun

May 15

May 15Fri
  1. Fidji SimoXAI score60

    ChatGPT adds a personal finance preview for U.S. Pro users

    AIChatGPT is previewing a personal finance experience for Pro users in the U.S., who can securely connect financial accounts and see where their money is going. Users can ask questions based on the information they choose to connect, and the author says this follows the similar health records connection feature.

  2. Intern Large ModelsOfficialAI score55

    Intern-S2-Preview: 35B Open Scientific Multimodal Model Released

    AIShanghai AI Laboratory's Intern Large Models introduces Intern-S2-Preview, a 35B scientific multimodal foundation model, and says it matches the trillion-scale Intern-S1-Pro on core scientific tasks. The post says it is the first open-source model with material crystal structure generation and strong general capabilities, with shared-weight MTP plus KL loss improving acceptance rate and speed. It is already supported by vLLM and SGLang, with weights on Hugging Face and ModelScope.

    Image from @intern_lm's post

May 7

May 7Thu
  1. Sam BowmanXAI score38

    Anthropic donates open-source alignment testing tool Petri to Meridian Labs

    AIAnthropic is donating Petri, its open-source interactive behavioral-evals tool for alignment testing, to Meridian Labs so development can continue independently. Working with Meridian, Anthropic has also released a major update improving the adaptability, realism, and depth of Petri's tests. Developers are invited to try the tool and contribute.

Apr 30

Apr 30Thu
  1. ARC PrizeOfficialAI score44

    GPT-5.5 and Opus 4.7 Fail ARC-AGI-3 Tasks Through Flawed World Models

    AIOpenAI's GPT-5.5 scored 0.43% and Anthropic's Opus 4.7 scored 0.18% on ARC-AGI-3, a set of 135 novel environments, according to ARC Prize's replay analysis of 160 runs. The analysis found three recurring failure modes: models perceived local action effects but failed to build global rules, mapped unfamiliar games onto known ones, and sometimes beat a level without learning the underlying mechanic. ARC Prize is open-sourcing its analysis package.

Apr 28

Apr 28Tue
  1. Andy JassyXAI score33

    Amazon Quick desktop app launches as personalized AI productivity assistant

    AIAmazon Quick's new desktop app connects to email, calendar, Slack, local files, and other apps to flag priority messages, summarize information, send communications, and create agents that handle tasks. The post says it becomes more personalized with use, and the author describes using it to treat their inbox more like an archive.

    Video from @ajassy's post
  2. Soumith ChintalaXAI score44

    Talkie: open-weight 13B LLM trained only on pre-1930 data

    AIResearchers including David Duvenaud and Alec Radford announced Talkie, an open-weight 13B-parameter LLM trained and finetuned on a newly curated dataset containing only pre-1930 text. Soumith Chintala shared it as a fun test of time for the model.

Apr 27

Apr 27Mon
  1. Xiaomi MiMo · new models on Hugging FaceOfficialAI score72

    Xiaomi releases MiMo-V2.5, an open omnimodal model with 1M context

    AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.

    Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.

  2. Mistral AI · new models on Hugging FaceOfficialAI score36

    Mistral Medium 3.5 EAGLE draft model released for speculative decoding on Hugging Face

    AIMistral AI has released mistralai/Mistral-Medium-3.5-128B-EAGLE, an EAGLE draft model for speculative decoding with the 128B dense Mistral Medium 3.5. The companion model, which the source says replaces Mistral Medium 3.1 and Magistral in Le Chat and Devstral 2 in Vibe, has a 256k context window, handles text and image input with text output, and is served with vLLM or SGLang using three speculative tokens. The model is released under a Modified MIT License that allows commercial use with exceptions for companies with large revenue.

Apr 26

Apr 26Sun
  1. Xiaomi MiMoOfficialAI score87

    Xiaomi releases open-source MiMo-V2.5-Pro for long-horizon agentic coding

    AIXiaomi released and open-sourced MiMo-V2.5-Pro, a 1.02T-parameter Mixture-of-Experts model with 42B active parameters and a 1M-token context window. The company reports gains in agentic tasks, software engineering, and long-horizon work, including a Rust SysY compiler task finished in 4.3 hours across 672 tool calls. Weights and tokenizer are on Hugging Face, and API pricing is unchanged.

    Why it matters: The release pairs a 1.02T-parameter open-weight model with long-horizon agent results and token-efficiency claims, useful for judging its fit in coding and agent workflows.

Apr 23

Apr 23Thu
  1. Apple · new models on Hugging FaceOfficialAI score40

    Apple releases CADD-Base-7B, a masked diffusion model for code generation

    AIApple has released CADD-Base-7B on Hugging Face, a 7B masked diffusion language model for code generation that uses Continuously Augmented Discrete Diffusion (CADD) to guide discrete denoising with a continuous flow-matching signal. The model loads through Transformers with trust_remote_code, and its diffusion_generate method supports CADD sampling modes "weighted" and "argmax" with alg options such as "entropy" and "maskgit_plus". The release builds on DiffuCoder and reuses Dream's modeling architecture and generation utilities.

  2. OpenAI Alignment Research BlogOfficialAI score44

    OpenAI Open-Sources Chain-of-Thought Monitorability Evaluation Datasets and Code

    AIOpenAI is releasing a subset of datasets, reference code, and the g-mean 2 metric for evaluating chain-of-thought monitorability. The release includes most datasets from its monitorability suite, while some evaluations relying on private or restricted data were omitted. The company says it will keep reporting monitorability results in future frontier reasoning model system cards.

Apr 22

Apr 22Wed

Apr 21

Apr 21Tue
  1. Xiaomi MiMoOfficialAI score67

    Xiaomi releases MiMo-V2.5, an open multimodal agent model with 1M context

    AIXiaomi released MiMo-V2.5, a 310B-parameter sparse MoE model with 15B active parameters that adds native visual and audio understanding. The model supports up to 1 million tokens of context, and its weights, tokenizer, and model card are available on Hugging Face. Xiaomi says it surpasses MiMo-V2-Pro on agentic performance and reports a Claw-Eval score of 62.3 on the general subset.

    Why it matters: The release pairs native visual and audio understanding with a 1M-token context window and open weights, a combination worth checking against your own multimodal workflows.

  2. NVIDIA AI DeveloperOfficialAI score29

    NVIDIA OpenShell v0.0.34 adds live sandbox policy updates and VM installs

    AINVIDIA's OpenShell v0.0.34 release lets users update sandbox policy without restarting the runtime. The update also adds install-vm, which installs the gateway and VM driver with new --driver-dir support, and sandbox get, which shows the active runtime policy. Supervisor seccomp improvements and HTTP normalization are included as well.

Apr 20

Apr 20Mon
  1. NVIDIA AI DeveloperOfficialAI score35

    OpenShell v0.0.33 adds hardened sandboxing and a standalone libkrun driver

    AINVIDIA released OpenShell v0.0.33, which adds seccomp and process-limit hardening, inference routing, and a standalone libkrun compute driver. The libkrun driver provides a lightweight VM backend as a second compute path for agents. The release also includes bug fixes, a docs refresh, and improved test stability.

Apr 17

Apr 17Fri
  1. OpenAI · new models on Hugging FaceOfficialAI score41

    OpenAI Releases Privacy Filter, an Open-Weight PII Detection Model on Hugging Face

    AIOpenAI released Privacy Filter, a bidirectional token-classification model that detects and masks personally identifiable information in text under the Apache 2.0 license. The model has 1.5B total parameters with 50M active, supports a 128,000-token context window, and can run in a web browser or on a laptop. Users can fine-tune it and adjust precision/recall tradeoffs through preset operating points.

Apr 16

Apr 16Thu