Skip to contentSkip to stories

Updated

#Open-source ecosystem

Showing low-relevance items too. Hide low-relevance items

Sep 12

Sep 12Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score58

    Shanghai AI Lab releases Intern-S2-397B, a 397B multimodal scientific model

    AIShanghai AI Lab's InternLM team released Intern-S2-397B, a multimodal foundation model for scientific intelligence and long-horizon agents. The model uses visual pre-training on raw scientific literature pages, multi-task reinforcement learning across more than 20 scientific domains, and agentic reinforcement learning in sandboxed environments.

  2. Mike KnoopXAI score46

    Mike Knoop urges keeping AI research open amid slowdown proposals

    AIMike Knoop says he sees a path to an ARC-AGI-4 benchmark focused on open-ended invention, which he calls the gating capability between zero-sum automation and positive-sum innovation. He argues that coordinated slowdown efforts would likely apply to everyone, including open-source work, and cites chain of thought and the transformer as inventions that grew out of open science research. He concludes the research frontier must stay open to keep humanity on a positive-sum path.

Sep 11

Sep 11Fri
  1. InferactOfficialAI score34

    vLLM adds day-0 support for DeepSeek v4.1 Flash across six NVIDIA GPUs

    AIInferact says vLLM now supports DeepSeek v4.1 Flash on day zero across H100, H200, B200, B300, GB200, and GB300 GPUs. SemiAnalysis independently verified the NVIDIA support, while the post notes AMD vLLM still does not work with the model. Serving recipes are available at recipes.vllm.ai.

  2. Meituan LongCatOfficialAI score12

    Meituan LongCat hosts Hugging Face team in Shanghai for open-source talks

    AIMeituan's LongCat team hosted the Hugging Face team at its Shanghai office to discuss future model development and open-source plans. The two sides also covered recent features and technical details of Hugging Face Transformers. The post frames the visit as a step toward wider collaboration across the open-source ecosystem.

    Image from @Meituan_LongCat's post
  3. Interconnects (Nathan Lambert)BlogAI score38

    Open-Source AI & Open Models Reading List Is Updated for Research and Policy Writing

    AINathan Lambert has compiled a reading list of open-model writing covering why labs release open weights, the open-versus-closed debate, and US-China competition, last updated 15 September 2026. The list includes pieces on open-model economics, safety and marginal-risk research, and recent Chinese releases such as Kimi K3 and GLM-5.2. It also cites lawmaker inquiries into Western companies' use of Chinese models.

  4. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score72

    Shanghai AI Lab releases Atria Dawn Preview, an agentic model built on GLM-5.2

    AIShanghai Artificial Intelligence Laboratory has released Atria Dawn Preview, an agentic model built on the 744B-parameter MoE GLM-5.2 foundation model, with a 256K context window. The release page reports benchmark results across search, coding, tool use, productivity, and cybersecurity, and describes text-only setup for Codex and Claude Code.

    Why it matters: The release page gives a full benchmark table against named rivals and setup steps for Codex and Claude Code, useful for anyone evaluating agentic models.

Sep 10

Sep 10Thu
  1. hardmaruXAI score52

    Sakana Fugu releases Fugu Max and Fugu Ultra v2 multi-agent orchestration models

    AISakana AI released Fugu Max and Fugu Ultra v2, multi-agent orchestration systems that route tasks across a pool of open-weights and specialized models. The source says Fugu Max delivers performance within striking distance of elite models at two to six times lower cost, while Fugu Ultra v2 outperforms Opus 5 and Fable 5 on Chartography and outperforms models costing three to five times more per token on DeepSWE.

    Image from @hardmaru's post
  2. Together AI BlogOfficialAI score52

    Together AI expands Fine-Tuning with live metrics, expert LoRA, and early stopping

    AITogether AI expanded its Fine-Tuning service with support for newer open-weight models, live metrics tracking, and finer training controls. Expert LoRA adapters can be applied to Mixture-of-Experts expert layers, and early stopping keeps the checkpoint with the best validation loss. Dataset previews, sample weights, pre-flight validation, and lower prices on selected models are also included.

  3. Ai2 · new models on Hugging FaceOfficialAI score34

    AstaBrief-8B-SFT: Ai2's 8B model for cited scientific research reports

    AIAi2 released AstaBrief-8B-SFT, an 8B intermediate supervised fine-tuning checkpoint built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report. On the ScholarQA-CS2 test set of 100 computer science questions, it scored an average of 83.7 versus 77.3 for base Qwen3-8B, with citation recall at 71.3 versus 64.6. The model is licensed under Apache 2.0 for research and educational use.

  4. Sebastian RaschkaXAI score19

    Single-GPU mixture-of-experts LLM trained from scratch in 8 days

    AIGiles Thomas extended the GPT-2-style code from Sebastian Raschka's "Build a Large Language Model (from Scratch)" into a 6-expert, 2-active mixture-of-experts model and trained it from scratch over 8 days. Raschka praised the project as interesting LLM work done on a single GPU.

  5. Cognition Blog (Devin, Windsurf)OfficialAI score22

    Cognition Welcomes Dioxus Team to Advance Open-Source Cross-Platform App Framework

    AICognition has welcomed Jonathan Kelley and the Dioxus team, whose framework Cognition used extensively to build and improve Devin's performance. Cognition plans to continue supporting Dioxus, Blitz, Taffy, and Subsecond while increasing investment in Dioxus-Native and Blitz. The Dioxus team will also work on Devin's virtual machine, computer use skills, and testing capabilities.

  6. RadixArkOfficialAI score22

    RadixArk publishes Miles cookbook for DeepSeek V4.1 Flash

    AIRadixArk has published a cookbook on its Miles documentation site covering how to run DeepSeek V4.1 Flash. The post itself contains only a link to the cookbook page, so no further details about features, figures, or setup steps are available.

  7. RadixArkOfficialAI score60

    Miles adds day-0 RL support for DeepSeek-V4.1-Flash

    AIRadixArk says Miles brings day-0 RL support to DeepSeek-V4.1-Flash, with SGLang providing inference support. The post says quantization-aware training mirrors SGLang's FP4/FP8 rounding, and that colocated training and rollout fit full-parameter RL on 16 GPUs. In a DAPO run over steps 0–80, per-token trainer–rollout KL stayed at 0.0012–0.0017 while reward rose from 0.51 to 0.78.

  8. LMSYS OrgOfficialAI score62

    SGLang adds day-0 inference and RL support for DeepSeek V4.1 Flash

    AISGLang and Miles ship day-0 inference and RL support for DeepSeek V4.1 Flash, with weights now available. The model is natively multimodal with 552B backbone parameters, 16B active during decode and 8B during prefill, and supports up to 1M context. V4.1 adds shared compressed KV across layers, a two-stage sparse indexer, and a 196B Engram lookup memory.

Sep 9

Sep 9Wed
  1. Fireworks AI BlogOfficialAI score58

    Fireworks AI outlines a staged path from closed APIs to owned specialized models

    AIFireworks AI describes a four-stage path for teams moving from renting closed frontier models to training their own, starting with API use and prompt, context, and harness engineering. The post uses the UIPad computer-use dataset to show that Kimi K3 ties GPT 5.6 Sol overall at 87.7 but wins three of four categories while costing about half as much, suggesting routing. After roughly three hours of training on the training split, the tuned Kimi K3 outperforms GPT 5.6 Sol on the held-out test set.

  2. TinkerOfficialAI score28

    Tinker and OpenResearch automate auditing of self-distillation methods

    AITinker says it and OpenResearch let agents test dozens of competing published post-training methods automatically, with compute cost forecast to within a dollar. The main post cites a grant-supported effort, while the quoted alphaXiv post says agents reproduced SDFT's continual learning benefits across Qwen3-8B and Qwen3-30B-A3B over multiple seeds.

  3. RadixArkOfficialAI score38

    RadixArk's Miles integrates SGLang for fast, aligned post-training rollouts

    AIRadixArk says its Miles framework natively supports SGLang for fast rollouts while keeping rollout and training aligned for reliable post-training at scale. The post thanks the community for contributions and feedback shaping Miles. A related post from @adarshxs describes Miles v0.1 running fully async agentic RL on a 744B MoE across 64 GB300 GPUs.

  4. Ai2 (Allen Institute for AI)OfficialAI score39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    AIGoodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.

Sep 8

Sep 8Tue
  1. Ian Johnson 🔬🤖XAI score22

    Latent Craft lets users explore a million book images via UMAP in browser

    AIIan Johnson introduced Latent Craft, a new way to explore large datasets with UMAP, letting users fly through and collect images from them. The demo covers all 1 million images explorable in the browser, drawn from a dataset of 1,080,814 public domain images, mostly from 19th-century books, shared on the Hugging Face Hub.

    Video from @enjalot's post
  2. Google Developers BlogOfficialAI score72

    Google releases ADK for Kotlin 1.0 for building production AI agents

    AIGoogle announced general availability of ADK for Kotlin 1.0, a Kotlin Multiplatform framework for building AI agents on servers and Android. Version 1.0 reaches feature parity with ADK 1.0 Core and adds Android extensions for on-device models, cloud Gemini via Firebase AI Logic, and persistent sessions and memory with Room and AppSearch. The post includes a server-side incident triage example using KSP-generated tools and skills, plus an Android financial assistant example with human confirmation for transfers.

    Why it matters: The post names the new Android and server-side capabilities and the code setup, helping Kotlin developers judge whether ADK fits their agent projects.

  3. InferactOfficialAI score42

    Inferact reports open models hit 130K tokens/GPU-sec on agentic workloads

    AIInferact says months of vLLM tuning for agentic workloads, validated on SemiAnalysis's AgentX benchmark, let open-source models reach up to 130K tokens per GPU-second. The company claims this is 106 times cheaper than Opus 5 API pricing. The work is described as part of a vLLM blog post covering architecture, framework, and runtime optimizations.

  4. Cohere · new models on Hugging FaceOfficialAI score38

    Cohere releases Tiny Aya Base 32K, a 3.35B multilingual model with 32K context

    AICohere Labs has released Tiny Aya Base 32K, an open-weights pretrained model with 3.35 billion parameters and a 32K context window. The model covers 70+ languages, including many lower-resourced ones, and is designed for downstream adaptation and long-context research. It is a base model that has not been instruction-tuned, and it is licensed under CC-BY-NC.

  5. BAAIOfficialAI score43

    FlagEval-Robo tests 12 open-weight embodied AI models across simulation and real robots

    AIBAAI introduces FlagEval-Robo, an open dual-track evaluation suite linking simulation with real-world execution. The team post-trained and stress-tested 12 leading open-weight embodied AI models under strictly aligned conditions. The post raises whether high benchmark scores reflect physical reality, though it does not yet report specific results.

    Image from @BAAIBeijing's post
  6. Daniel HanXAI score28

    Qwen3.8-27B GGUF becomes the most-liked GGUF on Hugging Face

    AIUnsloth's Qwen3.8-27B GGUF is now the most-liked GGUF ever on Hugging Face, with the post predicting it will soon enter the top 30 most-liked models overall. Unsloth reports it reached 10M downloads and 3.7K likes in 24 days, and credits the community, Hugging Face, and the Qwen team.

  7. Interconnects (Nathan Lambert)BlogAI score40

    Motif-3, GLM-5.3, Hy4-preview and open model licenses in latest roundup

    AIOpen model licenses are tightening at the Chinese frontier, with Zhipu's GLM-5.3 switching from MIT to a custom license requiring a security review for inference and fine-tuning providers with over $10 billion in annual revenue. Motif-3 ships under an MIT license with strong scores for its size, while Tencent's Hy4-preview is a competent model that currently overthinks. Western makers Google and Meta have moved to Apache 2.0.

  8. Mistral AIOfficialAI score62

    Mistral raises €3B Series D at over €21B valuation led by Samsung

    AIMistral announced a €3 billion Series D round at a post-money valuation of more than €21 billion, led by Samsung Electronics with co-leads Scaleup Europe Fund and PSG Equity. The company says the funding will expand frontier research, compute capacity, infrastructure, and international growth, and that it now operates in 20 countries with 125+ enterprise customers including Airbus, ASML, and HSBC.

    Why it matters: The round shows how a company frames sovereign, open-weight AI as a full stack spanning models, infrastructure, compute, and products, which is useful context for European enterprise AI strategy.

  9. NVIDIA · new models on Hugging FaceOfficialAI score46

    NVIDIA Releases NV-Reason-CT, a 3D Vision-Language Model for Chest and Abdominal CT

    AINVIDIA's NV-Reason-CT is a 3D vision-language model for CT image analysis that combines a native 3D vision encoder with a language model. It is designed for radiology report generation, question answering, and multi-step reasoning across chest and abdominal CT volumes. The model converts a 384×384×384-mm input into 13,824 visual tokens without spatial downsampling and is available on Hugging Face under the OpenMDW-1.1 License.

Sep 7

Sep 7Mon
  1. Tencent HyOfficialAI score44

    Tencent Hy4 preview upgraded to cut overthinking and token use

    AITencent Hunyuan says its Hy4 preview has been upgraded to reduce long thinking and over-verification on complex tasks, which users had flagged. The company reports the same task quality with fewer turns and lower input and output tokens, confirmed by benchmark and human evaluation. The upgrade is live for all users, and Tencent says it will keep iterating based on feedback.

Sep 6

Sep 6Sun
  1. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score35

    UltraData-Code-L2-Classifier scores files for algorithmic code selection

    AIOpenBMB released UltraData-Code-L2-Classifier, a suite of language-specific file-level scorers for 11 programming languages in UltraData-Code-L1. The L2 corpus selected with these scorers contains approximately 400B tokens and retains about 12.23% of L1 files, and a 10B-token test on a 1B model raised EvalPlus pass@1 by 7.80 points over L1 training.

  2. OpenBMB (MiniCPM) · new models on Hugging FaceOfficialAI score62

    OpenBMB releases MiniCPM5-2B, a 2B open-source model with open training data

    AIOpenBMB has released MiniCPM5-2B, a dense 2B Transformer built for on-device and resource-constrained deployment, with an average score of 53.9 in its comparison set. The release also opens the UltraData datasets behind it, including UltraX, UltraData-Code, UltraData-SFT-Agent-2609 and UltraData-RL-2609, and includes GGUF, MLX, GPTQ and DSpark variants for common runtimes.

    Why it matters: The release pairs a 2B model with open training datasets and reports per-benchmark comparisons against named same-size and larger models, letting readers check the claims directly.

Sep 5

Sep 5Sat
  1. AI at MetaOfficialAI score38

    AIRA₃ ensemble places 8th with gold-medal results in live competition

    AIMeta's AIRA₃ entered the live competition with an ensemble of models, and the 8th-ranked gold-medal entry combined GPT 5.5 (w/ OpenCode) and Claude 4.8 (w/ ClaudeCode). Post-hoc testing found Muse Spark 1.2 (w/ MuseCode) also reached gold-medal level, while Muse Spark 1.1 (w/ OpenCode) and GLM 5.2 (w/ OpenCode) reached silver-medal level, all graded on the same private test set.

    Image from @AIatMeta's post

Sep 4

Sep 4Fri
  1. Matei ZahariaXAI score46

    Qwen3.8-Flash-Next runs at 68.3 tok/s on a single RTX 5090

    AIA Berkeley Sky Lab researcher says stronger open models and new inference systems will make powerful local AI practical. The linked post reports Qwen3.8-Flash-Next running at 68.3 tok/s on a single RTX 5090 using an NVFP4 checkpoint, with 63GB host RAM and a 51GB n-gram table stored on NVMe at about 0.5% throughput cost.