Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

May 15

May 15Fri
  1. Fidji SimoXAI score60

    ChatGPT adds a personal finance preview for U.S. Pro users

    AIChatGPT is previewing a personal finance experience for Pro users in the U.S., who can securely connect financial accounts and see where their money is going. Users can ask questions based on the information they choose to connect, and the author says this follows the similar health records connection feature.

    Why it matters: The launch lets Pro users in the U.S. connect financial accounts to ChatGPT, showing how a product-level data integration is being extended beyond health records.

  2. Intern Large ModelsOfficialAI score55

    Intern-S2-Preview: 35B Open Scientific Multimodal Model Released

    AIShanghai AI Laboratory's Intern Large Models introduces Intern-S2-Preview, a 35B scientific multimodal foundation model, and says it matches the trillion-scale Intern-S1-Pro on core scientific tasks. The post says it is the first open-source model with material crystal structure generation and strong general capabilities, with shared-weight MTP plus KL loss improving acceptance rate and speed. It is already supported by vLLM and SGLang, with weights on Hugging Face and ModelScope.

    Image from @intern_lm's post

May 7

May 7Thu
  1. Sam BowmanXAI score38

    Anthropic donates open-source alignment testing tool Petri to Meridian Labs

    AIAnthropic is donating Petri, its open-source interactive behavioral-evals tool for alignment testing, to Meridian Labs so development can continue independently. Working with Meridian, Anthropic has also released a major update improving the adaptability, realism, and depth of Petri's tests. Developers are invited to try the tool and contribute.

Apr 30

Apr 30Thu
  1. ARC PrizeOfficialAI score44

    GPT-5.5 and Opus 4.7 Fail ARC-AGI-3 Tasks Through Flawed World Models

    AIOpenAI's GPT-5.5 scored 0.43% and Anthropic's Opus 4.7 scored 0.18% on ARC-AGI-3, a set of 135 novel environments, according to ARC Prize's replay analysis of 160 runs. The analysis found three recurring failure modes: models perceived local action effects but failed to build global rules, mapped unfamiliar games onto known ones, and sometimes beat a level without learning the underlying mechanic. ARC Prize is open-sourcing its analysis package.

Apr 28

Apr 28Tue
  1. Andy JassyXAI score33

    Amazon Quick desktop app launches as personalized AI productivity assistant

    AIAmazon Quick's new desktop app connects to email, calendar, Slack, local files, and other apps to flag priority messages, summarize information, send communications, and create agents that handle tasks. The post says it becomes more personalized with use, and the author describes using it to treat their inbox more like an archive.

    Video from @ajassy's post
  2. Soumith ChintalaXAI score44

    Talkie: open-weight 13B LLM trained only on pre-1930 data

    AIResearchers including David Duvenaud and Alec Radford announced Talkie, an open-weight 13B-parameter LLM trained and finetuned on a newly curated dataset containing only pre-1930 text. Soumith Chintala shared it as a fun test of time for the model.

Apr 27

Apr 27Mon
  1. Xiaomi MiMo · new models on Hugging FaceOfficialAI score72

    Xiaomi releases MiMo-V2.5, an open omnimodal model with 1M context

    AIXiaomi's MiMo-V2.5 is a native omnimodal model that understands text, image, video, and audio within one architecture. It is a sparse MoE with 310B total and 15B activated parameters, and supports up to 1M tokens of context. The repository also notes a config.json and tokenizer_config.json update that users who downloaded before commit 4da2748 should re-pull.

    Why it matters: The repository documents a 310B-parameter omnimodal MoE with a hybrid attention design, useful for comparing long-context efficiency against other open multimodal models.

  2. Mistral AI · new models on Hugging FaceOfficialAI score36

    Mistral Medium 3.5 EAGLE draft model released for speculative decoding on Hugging Face

    AIMistral AI has released mistralai/Mistral-Medium-3.5-128B-EAGLE, an EAGLE draft model for speculative decoding with the 128B dense Mistral Medium 3.5. The companion model, which the source says replaces Mistral Medium 3.1 and Magistral in Le Chat and Devstral 2 in Vibe, has a 256k context window, handles text and image input with text output, and is served with vLLM or SGLang using three speculative tokens. The model is released under a Modified MIT License that allows commercial use with exceptions for companies with large revenue.

Apr 26

Apr 26Sun
  1. Xiaomi MiMoOfficialAI score87

    Xiaomi releases open-source MiMo-V2.5-Pro for long-horizon agentic coding

    AIXiaomi released and open-sourced MiMo-V2.5-Pro, a 1.02T-parameter Mixture-of-Experts model with 42B active parameters and a 1M-token context window. The company reports gains in agentic tasks, software engineering, and long-horizon work, including a Rust SysY compiler task finished in 4.3 hours across 672 tool calls. Weights and tokenizer are on Hugging Face, and API pricing is unchanged.

    Why it matters: The release pairs a 1.02T-parameter open-weight model with long-horizon agent results and token-efficiency claims, useful for judging its fit in coding and agent workflows.

Apr 23

Apr 23Thu
  1. Apple · new models on Hugging FaceOfficialAI score40

    Apple releases CADD-Base-7B, a masked diffusion model for code generation

    AIApple has released CADD-Base-7B on Hugging Face, a 7B masked diffusion language model for code generation that uses Continuously Augmented Discrete Diffusion (CADD) to guide discrete denoising with a continuous flow-matching signal. The model loads through Transformers with trust_remote_code, and its diffusion_generate method supports CADD sampling modes "weighted" and "argmax" with alg options such as "entropy" and "maskgit_plus". The release builds on DiffuCoder and reuses Dream's modeling architecture and generation utilities.

  2. OpenAI Alignment Research BlogOfficialAI score44

    OpenAI Open-Sources Chain-of-Thought Monitorability Evaluation Datasets and Code

    AIOpenAI is releasing a subset of datasets, reference code, and the g-mean 2 metric for evaluating chain-of-thought monitorability. The release includes most datasets from its monitorability suite, while some evaluations relying on private or restricted data were omitted. The company says it will keep reporting monitorability results in future frontier reasoning model system cards.

Apr 22

Apr 22Wed

Apr 21

Apr 21Tue
  1. Xiaomi MiMoOfficialAI score67

    Xiaomi releases MiMo-V2.5, an open multimodal agent model with 1M context

    AIXiaomi released MiMo-V2.5, a 310B-parameter sparse MoE model with 15B active parameters that adds native visual and audio understanding. The model supports up to 1 million tokens of context, and its weights, tokenizer, and model card are available on Hugging Face. Xiaomi says it surpasses MiMo-V2-Pro on agentic performance and reports a Claw-Eval score of 62.3 on the general subset.

    Why it matters: The release pairs native visual and audio understanding with a 1M-token context window and open weights, a combination worth checking against your own multimodal workflows.

  2. NVIDIA AI DeveloperOfficialAI score29

    NVIDIA OpenShell v0.0.34 adds live sandbox policy updates and VM installs

    AINVIDIA's OpenShell v0.0.34 release lets users update sandbox policy without restarting the runtime. The update also adds install-vm, which installs the gateway and VM driver with new --driver-dir support, and sandbox get, which shows the active runtime policy. Supervisor seccomp improvements and HTTP normalization are included as well.

Apr 20

Apr 20Mon
  1. NVIDIA AI DeveloperOfficialAI score35

    OpenShell v0.0.33 adds hardened sandboxing and a standalone libkrun driver

    AINVIDIA released OpenShell v0.0.33, which adds seccomp and process-limit hardening, inference routing, and a standalone libkrun compute driver. The libkrun driver provides a lightweight VM backend as a second compute path for agents. The release also includes bug fixes, a docs refresh, and improved test stability.

Apr 17

Apr 17Fri
  1. OpenAI · new models on Hugging FaceOfficialAI score41

    OpenAI Releases Privacy Filter, an Open-Weight PII Detection Model on Hugging Face

    AIOpenAI released Privacy Filter, a bidirectional token-classification model that detects and masks personally identifiable information in text under the Apache 2.0 license. The model has 1.5B total parameters with 50M active, supports a 128,000-token context window, and can run in a web browser or on a laptop. Users can fine-tune it and adjust precision/recall tradeoffs through preset operating points.

Apr 16

Apr 16Thu

Apr 14

Apr 14Tue
  1. Moonshot AI (Kimi) · new models on Hugging FaceOfficialAI score78

    Moonshot AI releases open-source Kimi K2.6 multimodal agentic model

    AIMoonshot AI released Kimi K2.6, an open-source native multimodal agentic model with 1T total and 32B activated parameters and a 256K context length. The model card reports benchmark results against GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro across agentic, coding, reasoning, and vision tasks, and supports swarms of up to 300 sub-agents.

    Why it matters: The model card gives specific agent swarm scale, context length, and benchmark comparisons against several frontier models, useful for judging its coding and agent capabilities.

Apr 13

Apr 13Mon
  1. BAAIOfficialAI score40

    ClawKeeper v1.0 releases open-source security framework for OpenClaw AI agents

    AIBAAI announces ClawKeeper v1.0, an open-source security framework for OpenClaw AI agents, combining Skill-based command policies, Plugin-based runtime monitoring, and a Watcher system-level observer. The independent Watcher is designed to block high-risk operations such as prompt injections, key leaks, rogue commands, and remote code execution, even if the agent is compromised. The paper is available on arXiv and the project code is hosted on GitHub.

Apr 10

Apr 10Fri
  1. Awni HannunXAI score34

    Ollama and oMLX demo Qwen 3.5 and Gemma 4 on Apple silicon

    AIOllama and oMLX showed impressive demos running Qwen 3.5 and Gemma 4 locally on Apple silicon using MLX. Awni Hannun said the capabilities of local LLMs and their surrounding ecosystem have advanced considerably over the past couple of years.

Apr 9

Apr 9Thu
  1. Awni HannunXAI score24

    Running YOLO26 Locally on Apple Silicon with MLX

    AIA blog post and release show how to run YOLO26 locally using MLX. The linked background post describes YOLO26-MLX as a native Apple Silicon port without PyTorch or an external GPU, with up to 2.6x faster inference and up to 1.7x faster training.

    Image from @awnihannun's post

Apr 8

Apr 8Wed
  1. MiniMax · new models on Hugging FaceOfficialAI score78

    MiniMax releases open-weight MiniMax-M2.7 with agent and coding gains

    AIMiniMax has released MiniMax-M2.7 on Hugging Face, describing it as its first model to participate in its own evolution. The source reports 56.22% on SWE-Pro, 46.3% on Toolathon, and 62.7% on MM ClawBench, and says an internal version autonomously optimized a programming scaffold over 100+ rounds for a 30% performance improvement.

    Why it matters: The source ties its benchmark claims to a self-evolution process and a named comparison set, which helps readers weigh how the reported gains were achieved.

Apr 7

Apr 7Tue

Apr 6

Apr 6Mon
  1. Tri DaoXAI score32

    Fast Muon optimizer coming to Blackwell consumer GPUs

    AITri Dao says a fast Muon optimizer is coming to consumer cards, since its symmetric matmul kernels work once Blackwell consumer GPU mainloop support is in place. Background from @jcz42 reports Gram Newton-Schulz symmetric kernels now support RTX 5090, with 2x faster Newton-Schulz and 1.7x faster optimizer time on 15 layers of Gemma-4 E2B.

  2. Black Forest Labs · new models on Hugging FaceOfficialAI score41

    FLUX.2 Small Decoder offers faster, lower-VRAM drop-in replacement for FLUX.2 decoder

    AIBlack Forest Labs released FLUX.2 Small Decoder, a distilled VAE decoder that works as a drop-in replacement for the standard FLUX.2 decoder on Hugging Face. It decodes about 1.4x faster and uses about 1.4x less VRAM at decode time, with ~28M decoder parameters versus ~50M in the full decoder and minimal quality loss. It is available under the Apache 2.0 license and is compatible with FLUX.2-klein-4B, FLUX.2-klein-9B, FLUX.2-klein-9b-kv, and FLUX.2-dev.

Apr 3

Apr 3Fri
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score73

    Z.ai releases GLM-5.1, a flagship model for agentic engineering

    AIZ.ai has released GLM-5.1, its next-generation flagship model for agentic engineering, with stronger coding than GLM-5. The model is described as staying effective over longer agentic tasks, sustaining optimization over hundreds of rounds and thousands of tool calls. The release lists benchmark results including SWE-Bench Pro at 58.4 and Terminal-Bench 2.0 at 63.5, and local deployment is supported through SGLang, vLLM, xLLM, Transformers, and KTransformers.

    Why it matters: The release gives benchmark tables against several rival models, letting readers compare GLM-5.1's coding and agentic results with GLM-5 and frontier systems.

Apr 1

Apr 1Wed
  1. Jim FanXAI score62

    CaP-X open-sources agentic robotics toolkit, benchmark, and RL setup

    AIJim Fan announced the open-source release of CaP-X, an agentic robotics framework in which LLM-driven agents control robot arms and humanoids through perception and actuation APIs. The release includes CaP-Gym with 187 manipulation tasks across RoboSuite, LIBERO-PRO, and BEHAVIOR, and CaP-Bench, which evaluates 12 frontier LLMs and VLMs across 8 tiers. The post also reports that a 7B open-source model rose from 20% to 72% success after 50 RL training iterations, with synthesized programs transferring to real robots.

    Video from @DrJimFan's post

Mar 31

Mar 31Tue
  1. Mistral AI · new models on Hugging FaceOfficialAI score76

    Mistral Medium 3.5 releases as a 128B dense merged model with vision

    AIMistral AI released Mistral Medium 3.5, a dense 128B model with a 256k context window that handles instruction-following, reasoning, and coding in a single set of weights. It replaces Mistral Medium 3.1, Magistral, and Devstral 2, and reasoning effort is configurable per request. The model accepts text and image input and is released under a Modified MIT License that excludes companies with large revenue.

    Why it matters: The release merges instruction, reasoning, and coding into one 128B model with per-request reasoning control, giving developers one set of weights to compare against separate specialized models.

  2. Alibaba NLP (Tongyi) · new models on Hugging FaceOfficialAI score26

    LaSER-Qwen3-8B: Alibaba NLP's 8B dense retriever with latent reasoning released on Hugging Face

    AIAlibaba NLP released LaSER-Qwen3-8B, an 8B-parameter dense retriever built on Qwen/Qwen3-8B that internalizes explicit reasoning into latent space through continuous latent thinking tokens. The model scores 29.3 nDCG@10 on the BRIGHT benchmark, ahead of the rewrite-then-retrieve pipeline's 28.1, and carries a 4096-dimension embedding with an 8192-token maximum sequence length. It is licensed under MIT and adds about 1.7× latency over standard single-pass dense retrievers.

Mar 30

Mar 30Mon

Mar 26

Mar 26Thu
  1. Guillaume Lample @ NeurIPS 2024XAI score62

    Mistral releases Voxtral TTS text-to-speech model with open weights

    AIMistral has released Voxtral TTS, a text-to-speech model, alongside a blog post, a playground, a technical report, and model weights on Hugging Face. The post itself contains only links and no further details about the model's capabilities.

    Why it matters: The post links a playground, technical report, and open model weights, letting readers test and verify the release themselves.

  2. Guillaume Lample @ NeurIPS 2024XAI score62

    Mistral releases Voxtral TTS, its first open-weight speech model

    AIMistral's Voxtral TTS is its first speech model, presented as an open-weight text-to-speech model that reportedly delivers SOTA performance at significantly lower cost with very low latency. It combines autoregressive generation of semantic speech tokens with flow-matching for acoustic tokens, and a technical report on its training methodology is being released.

    Why it matters: The post names Voxtral TTS's architecture and a technical report, giving readers a concrete basis for comparing its speech generation method with other text-to-speech systems.

    Image from @GuillaumeLample's post
  3. Intern Large ModelsOfficialAI score44

    DataChef: RL framework auto-generates data recipes for LLM adaptation

    AIDataChef, an AI4AI framework, uses reinforcement learning to automatically generate optimal data recipes for adapting LLMs. Its DataChef-32B model, using an efficient proxy reward system, matches Gemini-3-Pro in recipe generation, with its recipes surpassing expert-curated ones on AIME'25 and ClimaQA benchmarks.

    Image from @intern_lm's post

Mar 24

Mar 24Tue
  1. ARC PrizeOfficialAI score70

    ARC Prize announces ARC-AGI-3, an interactive benchmark for frontier agents

    AIARC Prize has released ARC-AGI-3, a set of hundreds of interactive, turn-based environments with thousands of game-style levels, with no instructions or stated goals. Humans score 100% while frontier AI scores 0.51%. ARC Prize 2026 offers over $2 million in prizes for open-source solutions to ARC-AGI-2 and ARC-AGI-3.

    Why it matters: The benchmark's human versus frontier AI gap and its interactive design show how agent evaluation is shifting from instruction-following toward exploration and adaptation.

  2. Jim FanXAI score62

    Jim Fan warns that compromised LiteLLM package shows risks for AI agents

    AIJim Fan reposted a report that LiteLLM PyPI release 1.82.8 was compromised and contained a litellm_init.pth file that sends credentials to a remote server and self-replicates. He argues agents make this worse, since files like skills, configs, or PDFs read into context could spread malicious instructions. He concludes that agentic frameworks need guardrails and audited tooling.

Mar 22

Mar 22Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceOfficialAI score32

    PrismAudio Adds Reinforcement Learning to Video-to-Audio Generation with Chain-of-Thought Planning

    AIPrismAudio is a framework that integrates reinforcement learning into video-to-audio generation, using a Chain-of-Thought planning mechanism. It builds on ThinkSound by splitting single-step reasoning into four CoT modules for semantic, temporal, aesthetic, and spatial dimensions, each with targeted reward functions. Code, model weights, and datasets are released for research and educational use under the MIT License, and commercial use requires explicit author authorization.

Mar 20

Mar 20Fri
  1. Aman SangerXAI score55

    Cursor's Composer 2 is built on Kimi k2.5 base model with added training

    AIAman Sanger says Cursor's team evaluated many base models on perplexity-based evals and found Kimi k2.5 the strongest. Composer 2 was then built with continued pretraining and a 4x scale-up of high-compute RL, with Fireworks providing inference and RL samplers. The author admits Cursor should have named the Kimi base in its launch blog and says it will do so for the next model.