Skip to content

#Open source/Repo

Oct 8

TodayOct 8Thu8 items
  1. Xiaomi MiMoAI score44

    Xiaomi releases open-source MiMo-V2.5-ASR speech recognition model with dialect support

    Xiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.

  2. Databricks BlogAI score40

    Funke Brings Native HL7v2 Parsing to Databricks Lakehouse

    Databricks has released Funke, a Python and PySpark library and deployable pipeline that parses HL7v2 healthcare messages into native Spark types while preserving the full message hierarchy. It succeeds Smolder, the Scala data source Databricks open-sourced in 2021, and ingests through Auto Loader into Unity Catalog bronze and silver tables. Users can query segments, fields, components, and subcomponents directly with DataFrame or Spark SQL expressions.

  3. Claude Code · GitHub ReleasesAI score56

    Claude Code v2.1.295 adds hook failure blocking and gateway controls

    Claude Code v2.1.295 adds onFailure: "block" for command and HTTP hooks, so a hook that cannot start, times out, or exits unexpectedly blocks the action. The release also adds an optional models list for Claude apps gateway upstreams, plus upstream_request_id in the inference audit event, and fixes a range of MCP, plugin, and terminal issues.

  4. Codex · GitHub ReleasesAI score36

    Codex 0.162.0 adds managed worktree tools and clickable URLs in the TUI

    OpenAI's Codex 0.162.0 release adds tools for creating and listing managed Git worktrees from trusted local projects when the worktrees feature is enabled. The update also lets users pin tasks in the agent Command Center, copy transcript blocks with /copy, and make URLs clickable in approval headers, questions, and warnings, along with several Linux and Windows sandbox fixes.

  5. PyTorch BlogAI score62

    NVIDIA Dynamo adds session-level IDs to route and cache agentic inference

    NVIDIA Dynamo uses a unified session-level identifier to make its inference stack aware of agent sessions, subagents, and their KV cache across turns and tool calls. On SWE-bench, two TP4 MiniMax-M2 replicas on one 8xH100 node gained roughly 12-16% throughput from program-aware scheduling over KV-aware routing alone. The post also describes experimental shared-pool indexing and a proposed KvHint interface for session-aware cache policies in vLLM and SGLang.

    AIWhy it matters: The post explains how session identifiers let an inference stack track agent working sets, with measured throughput gains on SWE-bench and agentic RL rollouts.

  6. JetBrains AI BlogAI score62

    JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning

    JetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.

    AIWhy it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.

Oct 7

Oct 7Wed
  1. Google Developers BlogAI score62

    Google open-sources ML Drift, a cross-platform GPU engine for on-device AI

    Google's AI Edge Team open-sourced ML Drift under Apache 2.0, a GPU compute engine for on-device AI inference across OpenGL ES, OpenCL, Metal, and WebGPU. It serves as the core GPU acceleration engine within LiteRT and succeeds the legacy TFLite GPU delegate, which will no longer receive new features. The post cites benchmarks showing up to 40% lower frame latency in YouTube Shorts and up to 30% faster on-device performance in Adobe Lightroom and Photoshop.

    AIWhy it matters: The post explains how ML Drift unifies GPU shaders across platforms and replaces the TFLite GPU delegate, which matters for developers deploying on-device models.

  2. Claude Code · GitHub ReleasesAI score36

    Claude Code v2.1.293 adds Claude Haiku 5.5 and fixes dozens of bugs

    Claude Code v2.1.293 adds Claude Haiku 5.5 (claude-haiku-5-5), now the default Haiku model on the Anthropic API, with 1M context and pricing of $0.10/$0.50 per Mtok ($0.50/$2.50 for prompts over 100K). The release also adds agentType to the subagentStatusLine payload and isDeferred to $.tool.register, and fixes numerous issues including a memory leak in HTTP MCP connections.

  3. Hugging Face BlogAI score49

    Liquid AI Releases Open d1-3B and d1-omni-600M Edge Decision Models

    Liquid AI released two open-weight decision models, d1-3B and d1-omni-600M (experimental), built on its Liquid Foundation Models and available on Hugging Face. d1-3B scores 48.57 on the Decision Index 0.2.1, the highest among decision models under 10B parameters, and answers a question in 16 ms on an NVIDIA Jetson AGX Thor and under 50 ms on a Jetson Orin Nano. The models support text and images (d1-3B) or text with image or audio (d1-omni-600M).

  4. Microsoft ResearchAI score62

    Microsoft Research Asia releases Agent Lightning v1.0 for agentic RL with real harnesses

    Microsoft Research Asia has open-sourced Agent Lightning v1.0, a roughly 3,500-line agentic RL framework that trains the same agent harness used in deployment. In an end-to-end coding agent pipeline, Qwen3.5-9B rose from 41.8% to 56.4% Pass@1 on SWE-bench Verified using about 6,000 training samples. The framework runs agents as standard Kubernetes jobs without paid commercial sandbox services.

    AIWhy it matters: The source shows how training with the deployed agent harness avoids rebuilding agents, and reports concrete SWE-bench Verified gains from about 6,000 samples.

Oct 6

Oct 6Tue
  1. Liquid AI BlogAI score62

    Liquid AI releases open d1-3B and d1-omni-600M decision models for edge devices

    Liquid AI released two open-weight d1 decision models, d1-3B and d1-omni-600M, on Hugging Face. d1-3B scores 48.57 on the Decision Index v0.2.1 public split and answers a single question in 8 ms on an NVIDIA GeForce RTX 4090 and 50 ms on a Jetson Orin Nano. d1-omni-600M is an experimental checkpoint that handles text with images or audio and scores 15.95 on the same index.

    AIWhy it matters: The release pairs open-weight decision models with measured latency across Apple, NVIDIA, and Jetson hardware, showing how edge deployment changes what is practical.

  2. Gemini CLI · GitHub ReleasesAI score14

    Gemini CLI v0.63.0 released with retry indicator and auth loop fixes

    Gemini CLI v0.63.0 adds a retry progress indicator during connection recovery and fixes an infinite authentication loop caused by file contention, headless keyring issues, and supervisor state drops. The release also bounds tool output size and cleans up temporary directories when background shell execution exits, alongside fixes for MCP enablement config handling and stdin restoration after capability detection.

  3. Google DeepMindAI score67

    Google DeepMind releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    Google DeepMind has released EmbeddingGemma 2, an open 740 million parameter model that maps text, images, audio, and video into one embedding space. It is built on the Gemma 4 architecture under an Apache 2.0 license and supports an 8K token context window. The company reports a code benchmark gain from 68.76 to 78.68 on MTEB Code and says the model can run on-device with about 567MB of active RAM for the full multimodal version on a Google Pixel 11 Pro.

    AIWhy it matters: The release shows how a 740M-parameter embedding model can cover text, code, images, audio, and video on local hardware, with memory and storage figures to compare against other on-device options.

  4. Claude Code · GitHub ReleasesAI score40

    Claude Code v2.1.292 adds plugin marketplace flag and fixes security issues

    Claude Code v2.1.292 adds a --marketplace option to claude plugin install, which adds the marketplace if needed and then installs the plugin from it. The release also adds an effort parameter to the Agent tool and fixes several security issues, including permission prompts bypassed for network (UNC) file reads and a sandboxed read path that could return files outside approved access.

  5. Google DeepMind · The KeywordAI score72

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    Google DeepMind has released EmbeddingGemma 2, a 740-million-parameter embedding model that maps text, images, audio, and video into a shared space and runs on local hardware under an Apache 2.0 license. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions to 512, 256, or 128, and the model supports an 8K-token context window. The model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform availability coming soon.

    AIWhy it matters: The release shows how a 740M-parameter multimodal embedder runs locally with a 768-to-128 dimension truncation option, useful for judging on-device retrieval designs.

  6. Claude BlogAI score62

    Claude now works inside Google Docs, Sheets, and Slides in public beta

    Claude for Google Workspace is in public beta on all paid Claude plans, adding a sidebar to Google Docs, Sheets, and Slides. It can read the open file, edit text, build formulas, pivot tables, charts, and slides, and it asks for approval before changes unless the user chooses "Accept all edits." New Docs, Sheets, and Slides connectors in beta let Claude create and edit Google files from the chat, with access matching existing Google sharing permissions.

    AIWhy it matters: The source specifies how Claude edits Docs, Sheets, and Slides in place and where users keep control, which clarifies the practical workflow change.

Oct 5

Oct 5Mon
  1. Google Developers BlogAI score62

    EmbeddingGemma 2 releases multimodal embeddings with modular encoder loading

    Google released EmbeddingGemma 2, an open embedding model under the Apache 2.0 license that maps text, code, images, video, and audio into a shared 768-dimensional space. Developers can load a 270M-parameter text and code setup, or add vision and audio encoders up to a 740M-parameter full multimodal model. Matryoshka truncation to 256 or 128 dimensions reduces vector storage, with the guide noting quality losses on image, video, and speech retrieval at lower dimensions.

    AIWhy it matters: The guide gives concrete encoder sizes and dimension-storage tradeoffs, showing how to choose a configuration for text, code, image, video, and audio retrieval.

  2. Claude Code · GitHub ReleasesAI score31

    Claude Code v2.1.290 adds hook fixes, Deny button for sign-in, and new CLI commands

    Claude Code v2.1.290 adds serverToolUses to plugin turn.step results and agentId to tool.check hook events, so hooks can distinguish subagent permission checks. The release also adds a Deny button to the Claude apps gateway sign-in approval page, plus claude attach and claude logs accepting partial session names.

  3. GitHub Blog · AI & MLAI score63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    GitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    AIWhy it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

  4. Cloudflare Blog · AIAI score40

    Cloudflare Birthday Week 2026 unveils cf CLI, EmDash CMS, and post-quantum tools

    Cloudflare announced 46 products and updates during Birthday Week 2026, including the cf CLI for the entire Cloudflare API and EmDash, an open-source Astro-based serverless CMS whose plugins run in isolated Worker sandboxes. The company also said it plans to become a public certificate authority that issues free Merkle Tree Certificates for post-quantum authentication.

Oct 3

Oct 3Sat
  1. Claude Code · GitHub ReleasesAI score7

    Claude Code v2.1.289 fixes plugin, sandbox, and terminal rendering bugs

    Claude Code v2.1.289 fixes a series of bugs, including deny and ask rules being bypassed on nested parts of compound shell commands on managed machines. It also fixes terminal freezes on short code blocks with unclosed tags, Read deny rules not applying to files reached through symlinks in the IDE, and plugin panes that drew nothing for certain link formats. A change to claude auth status that may have increased sign-outs in VSCode was reverted.

  2. IndexTeam (Bilibili) · new models on Hugging FaceAI score22

    Index-Echo-S2ST-9B-FP4 released as NVFP4 quantized speech translation model

    IndexTeam released Index-Echo-S2ST-9B-FP4, an NVFP4 (W4A4) quantization of the Index-Echo-S2ST-9B speech-to-speech translation model, with only its text LLM backbone quantized. Perplexity rose from 3.8218 to 3.9650 (+3.75%) on a fixed corpus, while zh→en and en→zh outputs were semantically equivalent, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  3. IndexTeam (Bilibili) · new models on Hugging FaceAI score27

    Index-Echo-S2ST-2B FP4 Quantized Speech-to-Speech Translation Model Released on Hugging Face

    IndexTeam released Index-Echo-S2ST-2B-FP4, an NVFP4 (W4A4) quantized version of the Index-Echo-S2ST-2B speech-to-speech translation model, with only the text LLM backbone quantized and the audio components kept in BF16. On a fixed corpus, perplexity rose from 5.9332 to 6.4980 (+9.52%), while zh->en and en->zh generations matched the original. Full FP4 acceleration requires an NVIDIA Blackwell GPU, and the model loads via compressed-tensors in vLLM or transformers.

  4. IndexTeam (Bilibili) · new models on Hugging FaceAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-9B speech translation model

    IndexTeam published an NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-9B speech-to-text translation model, quantizing only the text LLM backbone while keeping the audio tower and other components in BF16. On an NVIDIA A100, perplexity rose from 3.4155 to 3.5113 (+2.81%), with zh->en and en->zh outputs semantically equivalent under greedy decoding. Full FP4 speedup requires an NVIDIA Blackwell GPU, while older GPUs get only memory reduction.

  5. IndexTeam (Bilibili) · new models on Hugging FaceAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-2B speech translation model

    IndexTeam has published an official NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-2B speech-to-text translation model on Hugging Face. Only the text LLM backbone is quantized, while the audio tower, connector, and speech-synthesis components remain in BF16. Perplexity rises 5.80%, from 4.8772 to 5.1599, on a fixed corpus, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  6. IndexTeam (Bilibili) · new models on Hugging FaceAI score22

    Index-Nailong-9B-FP4 NVFP4 quantized translation model released on Hugging Face

    IndexTeam released Index-Nailong-9B-FP4, an official NVFP4 (W4A4) quantization of the Index-Nailong-9B multilingual translation model, which covers 150 languages. In a validation on an NVIDIA A100 against the BF16 checkpoint, perplexity rose 3.10% (2.4339 to 2.5094), and zh-en and en-zh outputs were semantically equivalent. Full FP4 compute acceleration requires an NVIDIA Blackwell GPU, while older GPUs get memory savings only; the FP8 build is recommended for Hopper and Ampere.

  7. IndexTeam (Bilibili) · new models on Hugging FaceAI score23

    Index-Homura-9B-FP4 released with NVFP4 quantization for translation model

    IndexTeam released Index-Homura-9B-FP4, an official NVFP4 (W4A4) quantization of the Index-Homura-9B translation model from the Index-Translate family. On a fixed corpus, perplexity rose from 2.5386 in BF16 to 2.6245, a 3.38% increase, and zh->en generations matched the original. Full FP4 compute acceleration requires an NVIDIA Blackwell GPU, while older GPUs get only weight-only memory savings and the FP8 build is recommended for them.

  8. IndexTeam (Bilibili) · new models on Hugging FaceAI score29

    Index-Homura-2B-FP4 released as NVFP4 quantized translation model

    IndexTeam released Index-Homura-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Homura-2B multilingual translation model, which supports 150 languages. The quantized checkpoint shows a 5.73% perplexity increase over the BF16 original (3.5011 to 3.7017) on a fixed corpus, and its zh-en and en-zh outputs are semantically equivalent under greedy decoding. Full FP4 acceleration requires an NVIDIA Blackwell GPU, while the source recommends the FP8 build for Hopper and Ampere hardware.

Oct 2

Oct 2Fri
  1. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    Ai2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    AIWhy it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  2. Google ResearchAI score60

    Google's TEE-based federated learning system adds verifiable privacy guarantees

    Google announces a next-generation federated learning system that uses Trusted Execution Environments to provide verifiable, auditable data anonymization. The system publishes access policies to a public transparency log and is deployed in Gboard, which has launched English and Japanese next-word prediction models with stronger privacy guarantees and improved accuracy. Training time has also sped up significantly because computation moved to the server and is parallelized across many machines.

    AIWhy it matters: The post shows how Trusted Execution Environments make federated learning's privacy claims externally verifiable, rather than relying on trust in the server operator.

  3. Ai2 (Allen Institute for AI)AI score67

    Ai2 open-sources AstaBrief 8B, a fast open-weights scientific report model

    Ai2 released AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. In Asta's Generate a report feature, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The model is built on Qwen3-8B with supervised fine-tuning and DPO, and institutions can run its open weights on their own infrastructure.

    AIWhy it matters: The post explains the data filtering and one-pass generation choices behind a fast open-weights report model, showing what worked and what did not.

Oct 1

Oct 1Thu
  1. Cloudflare Blog · AIAI score58

    Cloudflare releases open-source Clef decision models and an RL fine-tuning service

    Cloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and launched a reinforcement learning fine-tuning service. In Cloudflare's tests, Clef classified a domain in 2.2s versus 4.7s for gpt-oss-120b, and the models are Jev-API compatible. The company is offering fine-tuning first through a forward-deployed engineering team, with a self-serve platform planned later.

Sep 30

Sep 30Wed
  1. Google Cloud · AI & Machine LearningAI score41

    Google Cloud Rolls Out Agent Substrate, GKE Agent Sandbox RL Tools in September

    Google Cloud introduced GKE Agent Substrate, an open-source execution runtime it says can run millions of sandboxes with 10x higher density than standard container runtimes. It also made GKE Agent Sandbox optimized for reinforcement learning generally available, alongside an orchestration SDK and native RL gym integrations. Google said GKE Pod snapshots can reduce AI inference start-up by as much as 89%, based on internal tests.

  2. Lovable BlogAI score47

    Lovable Discloses TanStack Start Vulnerability CVE-2026-102989 and Protects Hosted Apps

    Lovable's security team found a vulnerability (CVE-2026-102989) in TanStack Start, which allows attackers to run unwanted JavaScript in visitors' browsers via crafted links. Lovable reported it to TanStack and deployed firewall protections for hosted apps while a fix was prepared, and affected projects will be automatically updated on their next change or via the Security page. Lovable says it found no evidence of exploitation in reviewed logs, and apps hosted elsewhere must apply the upstream update themselves.

Sep 29

Sep 29Tue
  1. BAAI · new models on Hugging FaceAI score62

    BAAI releases AREX-2, a 27B agent model for self-improving long-horizon tasks

    BAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.

    AIWhy it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.

  2. Anthropic ResearchAI score80

    Anthropic says GLM-5.3 gives attackers cyber capabilities with weak safeguards

    Anthropic reports that Zhipu AI's GLM-5.3 can autonomously build end-to-end cyber exploits and is released without meaningful safeguards against misuse. In its simulated tests, attackers bypassed the model's safeguards 64% to 100% of the time using simple techniques, while the same attacks failed against safeguarded Claude models. Anthropic also cites an NIST CAISI assessment calling GLM-5.3 the most cyber-capable open-weight model released to date.

    AIWhy it matters: The report shows how open-weight safeguards fail under simple bypasses, offering concrete test figures for judging misuse risk in released models.

Sep 27

Sep 27Sun
  1. Xiaomi MiMo · new models on Hugging FaceAI score44

    Xiaomi releases MiMo-V2.6-Flash-MOPD, an upgraded MoE model with 1M context

    Xiaomi has released MiMo-V2.6-Flash-MOPD on Hugging Face, an upgrade of the MiMo-V2.6-Flash-RL checkpoint that fuses several domain-specialized teachers into one model. The sparse MoE model has 309B total and 15B activated parameters, a 1M-token context length, and supports text, image, video, and audio inputs. The checkpoint targets tool-call repetition, a failure mode where the model repeatedly issues the same or similar tool calls without making progress.