Skip to content

#Open source/Repo

Oct 8

TodayOct 8Thu1 item

Oct 7

Oct 7Wed
  1. a16z NewsAI score46

    a16z backs Preference Model, which builds RL environments for training AI models

    Preference Model is open-sourcing Karotte, the framework it uses to build reinforcement learning environments that resist reward hacking, including defenses like killing stray processes before grading and rejecting grader-crashing files. The framework has been hardened through more than a million evaluation runs and controlled red-teaming. The company focuses on machine learning engineering tasks for leading labs, and a16z says it is partnering with Preference Model and its founders, Jennifer Zhou and Ning Cao.

  2. Semafor · TechnologyAI score56

    Reflection AI and Mistral launch open models to challenge China's lead

    Reflection AI and Mistral each unveiled new open-source models this week, aiming to beat other Western open models, though they trail top Chinese and closed systems on prominent benchmarks. Reflection CEO Misha Laskin says the target is regulated industries and governments that cannot or will not use Chinese models. The outcome depends on whether businesses and agencies accept less advanced models for some tasks in exchange for lower cost and more control.

  3. Testing CatalogAI score47

    Daily AI brief covers Mistral Large 4, Google, OpenAI, and Anthropic updates

    Mistral released Mistral Large 4 "Le Chonk", a 1T-parameter (49B active) multimodal model, with open weights planned in about three weeks. Google rolled out Nano Banana 2.1 across Gemini, AI Studio, and the Gemini API, and released EmbeddingGemma 2, a 740M-parameter open multimodal embedding model under Apache 2.0. OpenAI launched the Decisions API in beta with gpt-6-luna, returning typed answers 10x faster than the Responses API.

Oct 6

Oct 6Tue
  1. METRAI score40

    This demonstration used a JavaScript injection via MathJax rendering, which an agent could trigger from anywhere in its output, including its reasoning. For clarity, we haven't seen agents exploit this in our evaluations, and Inspect was patched within a day of our report.

    This demonstration used a JavaScript injection via MathJax rendering, which an agent could trigger from anywhere in its output, including its reasoning. For clarity, we haven't seen agents exploit this in our evaluations, and Inspect was patched within a day of our report.

Oct 5

Oct 5Mon
  1. GeekPark (极客公园)AI score38

    OpenAI Launches 28-Day Codex and ChatGPT Work Improvement Plan, Adds Visual Ads in ChatGPT

    OpenAI says it will ship one meaningful Codex and Work improvement each day for 28 days starting October 5, or else offer a "reset" without specifying what that reset covers. The company also plans to test visual ads in ChatGPT image generation in the U.S. starting in late October, with ads kept separate from generated images and not affecting answers.

  2. Alex HeathAI score52

    Reflection's founders discuss building a DeepSeek of the West with Beam

    Reflection is set to release Beam, its first open-weight AI model, aiming to become a Western counterpart to DeepSeek. The source says Beam is trained from scratch for coding, reasoning, and AI agents, with benchmarks placing it alongside the strongest open models and more efficient token economics. Reflection has raised $4.6 billion from investors including Nvidia, Sequoia, and Lightspeed, and the interview covers its monetization plans for open-weight models.

Oct 3

Oct 3Sat
  1. SemiAnalysisAI score34

    BIG MOMENT FOR AMD🚨THIS WEEK: AMD HAS REACHED ABOVE 90% PARITY ON UPSTREAM vLLM GATING TEST GROUPS! This was after months of hard work from AMD maintainers, specifically Andreas, and also vLLM CI Lead Kevin, along with SemiAnalysis getting vLLM enough CI AMD GPUs after months of upstream maintainer complaints about not having enough AMD GPUs on vLLM and AMD CI fleet-wide stability issues. (1/3)🧵

    BIG MOMENT FOR AMD🚨THIS WEEK: AMD HAS REACHED ABOVE 90% PARITY ON UPSTREAM vLLM GATING TEST GROUPS! This was after months of hard work from AMD maintainers, specifically Andreas, and also vLLM CI Lead Kevin, along with SemiAnalysis getting vLLM enough CI AMD GPUs after months of upstream maintainer complaints about not having enough AMD GPUs on vLLM and AMD CI fleet-wide stability issues. (1/3)🧵

Oct 2

Oct 2Fri
  1. François CholletAI score28

    Keras community call outlines pluggable backends and KerasHub updates

    Keras is moving to a pluggable backend design, with MLX and PaddlePaddle backends upcoming as add-on libraries. The team is reducing the operations needed to ship new backends and streamlining unit testing so a single harness can test all ops, such as casting consistency. KerasHub also gains many new models and is shifting its preprocessing from tf-text to PyGrain.

Oct 1

Oct 1Thu
  1. GoodfireAI score58

    Goodfire Says AI Biosecurity Risks Are Next After Cybersecurity Risks

    Goodfire says AI cybersecurity risks are already here and that biosecurity risks are next, as models improve at biology. The post presents this as both an opportunity for science and medicine and a reason for stronger security. It quotes Demis Hassabis announcing SynthID for biology, a watermarking approach for AI-generated proteins, published in Nature with SynthID Bio tools open sourced.

Sep 30

Sep 30Wed
  1. Thomas WolfAI score31

    ESM-2 came out in 2022. It's still downloaded hundreds of thousands of times a month. That only works because someone keeps maintaining the software underneath it. @huggingface 🤝 @os4science are teaming up to find those libraries and back the people behind them 🧬 https://os4science.org/news/hugging-face-open-source-for-science-fund/

    ESM-2 came out in 2022. It's still downloaded hundreds of thousands of times a month. That only works because someone keeps maintaining the software underneath it. @huggingface 🤝 @os4science are teaming up to find those libraries and back the people behind them 🧬 https://os4science.org/news/hugging-face-open-source-for-science-fund/

  2. Lovable BlogAI score47

    Lovable Discloses TanStack Start Vulnerability CVE-2026-102989 and Protects Hosted Apps

    Lovable's security team found a vulnerability (CVE-2026-102989) in TanStack Start, which allows attackers to run unwanted JavaScript in visitors' browsers via crafted links. Lovable reported it to TanStack and deployed firewall protections for hosted apps while a fix was prepared, and affected projects will be automatically updated on their next change or via the Security page. Lovable says it found no evidence of exploitation in reviewed logs, and apps hosted elsewhere must apply the upstream update themselves.

  3. Daniel HanAI score7

    We're co-hosting an open source party with @HuggingFace! 🤗🦥 Come join us, we'll be handing out @UnslothAI merch, showcasing our upcoming features and more! It'll be a super fun night with lots of demos, DJs, food and more. Join: http://luma.com/OpenTogether Use code: NOSLOTHSHERE to get in.

    We're co-hosting an open source party with @HuggingFace! 🤗🦥 Come join us, we'll be handing out @UnslothAI merch, showcasing our upcoming features and more! It'll be a super fun night with lots of demos, DJs, food and more. Join: http://luma.com/OpenTogether Use code: NOSLOTHSHERE to get in.

Sep 29

Sep 29Tue
  1. PerplexityAI score10

    Our Secure Intelligence Institute works with researchers at Stanford, CMU, Duke, Columbia, Ohio State, and UVA. We’re also working with NVIDIA and the Open Secure AI Alliance, and sharing tools and findings so others can strengthen their own systems. https://blogs.nvidia.com/blog/open-secure-ai-alliance/?ncid=so-twit-957725

    Our Secure Intelligence Institute works with researchers at Stanford, CMU, Duke, Columbia, Ohio State, and UVA. We’re also working with NVIDIA and the Open Secure AI Alliance, and sharing tools and findings so others can strengthen their own systems. https://blogs.nvidia.com/blog/open-secure-ai-alliance/?ncid=so-twit-957725

  2. Anthropic ResearchAI score80

    Anthropic says GLM-5.3 gives attackers cyber capabilities with weak safeguards

    Anthropic reports that Zhipu AI's GLM-5.3 can autonomously build end-to-end cyber exploits and is released without meaningful safeguards against misuse. In its simulated tests, attackers bypassed the model's safeguards 64% to 100% of the time using simple techniques, while the same attacks failed against safeguarded Claude models. Anthropic also cites an NIST CAISI assessment calling GLM-5.3 the most cyber-capable open-weight model released to date.

    AIWhy it matters: The report shows how open-weight safeguards fail under simple bypasses, offering concrete test figures for judging misuse risk in released models.

Sep 28

Sep 28Mon
  1. Khazix (数字生命卡兹克)AI score38

    Khazix open-sources AIHOT, a million-MAU AI news site, on GitHub

    数字生命卡兹克 announced that AIHOT, an AI hotspot news site with about one million monthly active users, is now open source on GitHub. The release includes the collection pipeline, curation scoring, clustering mechanism, and the production prompts, aiming to let others build vertical versions for industries such as gaming, law, HR, and finance.

Sep 25

Sep 25Fri
  1. OpenClawAI score46

    Today Microsoft announced Autopilot, an always on agent built on OpenClaw The best part of this collaboration is how much @OmarShahine and others at Microsoft have contributed BACK to OpenClaw Read all about the contributions Omar and his team made here: https://openclaw.ai/blog/microsoft-autopilot-openclaw

    Today Microsoft announced Autopilot, an always on agent built on OpenClaw The best part of this collaboration is how much @OmarShahine and others at Microsoft have contributed BACK to OpenClaw Read all about the contributions Omar and his team made here: https://openclaw.ai/blog/microsoft-autopilot-openclaw

Sep 24

Sep 24Thu
  1. NVIDIAAI score42

    With @GoogleDeepMind, @emblebi, and research partners, we’re making AI-predicted protein complex structures for 2,800+ viruses openly available. This gives scientists a head start in preparing for potential outbreaks.

    With @GoogleDeepMind, @emblebi, and research partners, we’re making AI-predicted protein complex structures for 2,800+ viruses openly available. This gives scientists a head start in preparing for potential outbreaks.

  2. inclusionAI (Ant Ling) · new models on Hugging FaceAI score22

    inclusionAI Publishes Training-Content Summaries for Ling and Ring Models

    inclusionAI has published public training-content summaries on Hugging Face for its Ling and Ring model versions, including Ling-2.0, Ling-2.5, Ling-2.6-1T, Ling-3.0, Ring-2.0, Ring-2.5-1T, and Ring-2.6-1T. The documents, organized under the template associated with Article 53(1)(d) of Regulation (EU) 2024/1689, contain documentation only, not model weights or training datasets. Each summary covers only the model versions it names.

Sep 23

Sep 23Wed
  1. Mike KnoopAI score57

    Tufa Labs reaches 83.06% on ARC-AGI-2, 2% short of the grand prize

    Mike Knoop says the top ARC Prize 2026 ARC-AGI-2 score of 83.06% by Tufa Labs is only 2% short of the 85% grand prize threshold. The challenge runs under strict Kaggle compute limits with no internet access, and the winning solution is set to be open sourced. The image shows the leaderboard with RabbitHole at 76.94%, nvbanana at 74.17%, Yi-Chia Chen at 55.14%, and Kha Vo at 37.50%.

Sep 22

Sep 22Tue
  1. Ant LingAI score34

    Thanks @ValsAI for the high-caliber eval! The term "flash" is a bit "misleading" now. With 124B total size and 5.1B activation, Ling-3.0-flash-fin is a "flash lite" with high intelligence density. Enjoy the free API while it last. We also have fp4 quant to be used on local AI 😛

    Thanks @ValsAI for the high-caliber eval! The term "flash" is a bit "misleading" now. With 124B total size and 5.1B activation, Ling-3.0-flash-fin is a "flash lite" with high intelligence density. Enjoy the free API while it last. We also have fp4 quant to be used on local AI 😛

Sep 20

Sep 20Sun

Sep 16

Sep 16Wed

Sep 15

Sep 15Tue
  1. Sundar PichaiAI score42

    Google outlines AI for science, weather, languages, and economic research

    Google says it is focusing AI efforts on health, disaster and weather resilience, learning, and economic opportunity. Recent examples include AlphaGenome Atlas, which maps all 9B possible single-letter genetic changes across the human genome and is openly available to researchers, and WeatherNext 3, described as its most accurate and capable global weather AI model to date. The post also cites AI & Economy ATLAS, an open-access look at global AI usage, and says its translation services now cover nearly 300 languages spoken by 7B people.

Sep 10

Sep 10Thu
  1. DeepSeekAI score37

    🌐 Supporting open source. Expanding deployment options. We’ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options. Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let’s talk. 🔹 Model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash 🔹 Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf 6/6

    🌐 Supporting open source. Expanding deployment options. We’ll work closely with the open-source community on V4.1-Flash inference support and explore more deployment options. Planning a large-scale deployment with 2,000 GPUs + a storage cluster? Let’s talk. 🔹 Model: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash 🔹 Paper: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf 6/6

Sep 8

Sep 8Tue
  1. Unsloth AIAI score25

    Qwen3.8-27B Unsloth GGUF is now the #1 most-liked GGUF of all time! The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥 GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF Guide: https://unsloth.ai/docs/models/qwen3.8

    Qwen3.8-27B Unsloth GGUF is now the #1 most-liked GGUF of all time! The model hit 10M downloads and 3.7K likes in just 24 days on Hugging Face - all thanks to you. 🤗🦥 GGUF: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF Guide: https://unsloth.ai/docs/models/qwen3.8

Sep 3

Sep 3Thu

Aug 24

Aug 24Mon
  1. Thinking MachinesAI score34

    Today, we are launching Tinker grants of up to $50,000 in credits for safety research on open-weight models. We share some project ideas that excite us below; if you’re working on a safety project that could be accelerated by additional Tinker credits, we want to hear from you!

    Today, we are launching Tinker grants of up to $50,000 in credits for safety research on open-weight models. We share some project ideas that excite us below; if you’re working on a safety project that could be accelerated by additional Tinker credits, we want to hear from you!

Aug 14

Aug 14Fri
  1. Z.aiAI score62

    Z.ai previews GLM-5.3 cyber model with staged release and OpenVuln initiative

    Z.ai says GLM-5.3 is its most capable model for cybersecurity tasks, with CyberGym at 84.5% versus 77.2% for GLM-5.2 and ExploitBench at 54.4% versus 24.4%. Access will begin with selected security partners in controlled settings, followed by broader access and API availability, with full open weights to be published after safety evaluations are complete. The company also launched the OpenVuln initiative to help open-source maintainers audit projects and coordinate disclosure.

Aug 11

Aug 11Tue
  1. Rowan CheungAI score62

    Meta opens weights for Muse Glimmer 30B model, Muse Spark 1.2 to follow

    Meta announced it is opening the weights for Muse Glimmer, a 30B parameter dense model that can run locally. Muse Spark 1.2, described as its latest foundation model, will have its weights released soon. The author's interview with Mark Zuckerberg quotes him saying Llama 4 fell short of the trajectory he wanted and that the lab was rebuilt.

Aug 10

Aug 10Mon