Skip to contentSkip to stories

Updated

#Open-source ecosystem

Showing low-relevance items too. Hide low-relevance items

Sep 22

Sep 22Tue
  1. Boris ChernyXAI score42

    Boris Cherny Uses Opus 5.5 to Formally Verify Claude Agent SDK

    AIBoris Cherny used Opus 5.5 to formally verify the Claude Agent SDK with Lean, and short prompts produced 16 PRs fixing bugs and race conditions. He also combines Lean and TLA+ to find issues in data flow, concurrency, and state management, and says Claude is strong in both languages even though he does not know them well.

    Video from @bcherny's post
  2. whXAI score34

    MiMo-V2.6 paper details data and RL results for open model

    AIThe MiMo-V2.6 paper thread reports on the newest open model, which also streams its RL run, focusing on data and RL experimental results rather than architecture. The Pro model reportedly rose from 58.41 to 72.57 on DeepSWE after RL, with the top published DeepSWE score cited at 74.

    Image from @nrehiew_'s post
  3. Comfy BlogOfficialAI score42

    ComfyUI Speeds Up MiniMax H3 Video VAE Encoding and Decoding

    AIComfyUI's update makes the MiniMax H3 video VAE encode up to about 2.2x faster and decode 1.4-2.7x faster, cutting a 1344x768, 129-frame round trip on an RTX 5090 from 24.3 to 12.7 seconds. The gains come from a fused encoder kernel enabled by default, fp16 accumulation support in a custom convolution, and an int8 decoder, and the source says the changes are visually lossless to the eye. Users need ComfyUI v0.36.0 or above, and the int8 VAE file is a drop-in replacement for the standard one.

  4. Daniel HanXAI score42

    Qwen-Image-2.1 runs locally in Unsloth Desktop via INT8, FP8, GGUF

    AIDaniel Han says Qwen-Image-2.1 works in Unsloth Desktop through INT8, FP8, and GGUF builds, with Unsloth also releasing dynamic GGUFs for it. Pinned RAM offloading lets INT8 and FP8 fit under 6–8GB of VRAM while remaining relatively fast. The linked Unsloth post says the 7B model runs on 12GB VRAM and performs on par with Nano Banana 2.0.

  5. Unsloth AIOfficialAI score70

    Qwen-Image-2.1 runs locally on 12GB VRAM using Unsloth GGUFs

    AIUnsloth says the 7B Qwen-Image-2.1 text-to-image and editing model can run locally on 12GB VRAM using its GGUF builds. It also states that the model performs on par with Nano Banana 2.0, and that Dynamic FP8 can run on 6GB of VRAM via offloading for higher quality. The image lists int8 at 7.26 GB with mean LPIPS 0.064 and fp8 at 7.12 GB with mean LPIPS 0.112, and says int8 is the default.

    Why it matters: The post gives concrete local-run settings, VRAM figures, and GGUF and FP8 options, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  6. Sebastian RaschkaXAI score62

    Xiaomi MiMo-V2.6-Pro tops open-weight benchmarks with simple attention design

    AIXiaomi's MiMo-V2.6-Pro ranks first among open-weight models on the Artificial Analysis Intelligence Index with a score of 46. The author attributes its standing mainly to a training data and post-training recipe that increased agent tasks and used an agentic grader for rewards, rather than its plain Grouped Query Attention and Sliding Window Attention design with a 128-token window.

    Image from @rasbt's post
  7. Black Forest Labs · new models on Hugging FaceOfficialAI score60

    Black Forest Labs releases FLUX 3 Action base weights for robot adaptation

    AIBlack Forest Labs has released flux-3-action-base, an open-weights 7B world action model that takes camera frames, robot state, and a text instruction to output the next action chunk. The release is an adaptation component rather than a complete robot policy, and new embodiments require their own action heads. The source says the weights are paired with shared video VAE and Qwen3-VL-4B-Instruct text encoders and is governed by the FLUX Kommunity License v.1.0.

    Why it matters: The source separates the adaptation base from full robot policies and states the shared encoders and new-embodiment requirements, which clarifies what developers must still build for their robots.

  8. OpenBMBOfficialAI score20

    OpenBMB praises MiniCPM5-2B workers in multi-agent invoice reconciliation

    AIOpenBMB thanked a developer for testing MiniCPM5-2B as a worker in a multi-agent workflow handling invoice matching, short payments, duplicate references, and disputes through tool calls. The background post says GPT-6 Astra coordinated the MiniCPM5-2B workers, verifying 32 synthetic invoices in 67.8 seconds with 232 executed tool calls. The demo does not move money.

  9. X.PINXAI score46

    Moonshot's Kimi K3 now available on Amazon Bedrock

    AIMoonshot's Kimi K3 is now available on Amazon Bedrock, with its license requiring a paid agreement for model-hosting businesses and affiliates above $20M in annual revenue. AWS says customer data stays within its cloud, is not shared with Moonshot or used for training, and inference requests have zero data retention. Neither company disclosed financial terms.

    Image from @thexpin's post

Sep 21

Sep 21Mon
  1. StepFunOfficialAI score58

    StepFun's Step 5 Preview scores 44 on Intelligence Index at lower cost

    AIStepFun's Step 5 Preview scores 44 on the Artificial Analysis Intelligence Index at about $0.72 per task, matching Kimi K3 (max) at roughly 2.8x lower cost. The source reports strong reasoning results, including 46% on Humanity's Last Exam, but places it behind Qwen3.8 Max and GLM-5.3 (max) on agentic evaluations. Open weights are planned for October 15.

  2. Tencent HyOfficialAI score67

    Tencent Hy4 preview compressed to 214 GiB with mixed-precision quantization

    AITencent Hunyuan says it shrank the 770B-parameter Hy4 preview from roughly 1.5TB to 214 GiB while keeping the parameter count unchanged. The quoted Zhihu post by a Tencent Hunyuan quantization team member describes the method: a 1.25-bit sparse ternary encoding, mixed precision across expert layers, and STQ1_0 CUDA kernels in llama.cpp. The author reports nearly unchanged MRCR retrieval and a small decline in math.

    Why it matters: The quoted Zhihu post explains how Hy4 preview's weights were quantized and kept usable at inference, a concrete engineering case for compressing large MoE models.

  3. vLLM BlogOfficialAI score60

    vllm-metal brings concurrent vLLM serving to Apple Silicon Macs

    AIvllm-metal ports vLLM's scheduler, paged KV cache, and OpenAI-compatible server to Apple Silicon, with MLX and Metal handling execution. The v0.28.0 release added batched MTP, GGUF and hybrid-model support, and faster prefill on M5, and v0.29.0 is installable through Homebrew.

    Why it matters: The post explains how vllm-metal packs requests and pages KV cache on Apple Silicon, with benchmarks showing where concurrent serving gains and tradeoffs appear.

  4. Xiaomi MiMoOfficialAI score13

    Xiaomi MiMo-V2.6-Pro ranks eighth on Design Arena

    AIXiaomi's MiMo-V2.6-Pro reached eighth overall and third among open-weight models on Design Arena with an Elo of 1338. That is a 54-point gain and 22-position climb from MiMo-V2.5-Pro, and the model also placed fourth in Website and sixth in Agentic Frontend Development, per Design Arena.

  5. Xiaomi MiMoOfficialAI score67

    Xiaomi MiMo open-sources Pro, Flash, and a 9B distilled model

    AIXiaomi MiMo announced open-source releases of Pro and Flash, the MiMo-V2.6-Distill-Qwen-9B model, a technical report, over 7K RL task environments, an end-to-end RL framework, and composable mini-harnesses. The attached table shows MiMo-V2.6-Distill-Qwen-9B after SFT and after RL compared with Qwen3.5-9B, with RL scores higher on most listed benchmarks, such as SWE-bench Verified at 66.2 versus 60.0.

    Why it matters: The table compares a 9B distilled model against Qwen3.5-9B on coding, cyber, and agent benchmarks, showing how the reinforcement learning stage changes results.

    Image from @XiaomiMiMo's post
  6. Xiaomi MiMoOfficialAI score78

    Xiaomi releases open-weight MiMo-V2.6 Pro and Flash omnimodal models

    AIXiaomi MiMo has launched MiMo-V2.6 Pro and Flash, two omnimodal models with open model weights, a technical report, RL environments, and training code. The post says Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks and scores 46 on the Artificial Analysis Intelligence Index, the highest among open-source models. A benchmark table compares Pro and Flash with MiMo-V2.5 Pro and frontier models across code agent, general agent, cybersecurity, and visual agent tests.

    Why it matters: The source pairs open-weight release details with a benchmark table against Claude Opus 5 and GPT-5.6 Sol, letting readers compare Pro and Flash across agent tasks.

    Image from @XiaomiMiMo's post
  7. RadixArkOfficialAI score25

    RadixArk's Miles adds async rollout buffer as swappable RL primitive

    AIRadixArk says its Miles framework uses an async rollout buffer that can change which sample groups reach training and which prompts get retried, while reusing the rollout worker and trainer. The post argues that stable, granular extension points let contributors modify one part of an RL system without disrupting its neighbors.

  8. Xiaomi MiMo · new models on Hugging FaceOfficialAI score50

    Xiaomi MiMo Releases MiMo-V2.6-Distill-Qwen-9B SFT Checkpoint on Hugging Face

    AIXiaomi MiMo released MiMo-V2.6-Distill-Qwen-9B, a 9B agentic model made by supervised fine-tuning Qwen3.5-9B on MiMo-generated data, as an open starting point for agentic reinforcement learning research. It scored 61.1 on SWE Verified, versus 60.0 for Qwen3.5-9B, and 44.6 on SWE Pro, versus 32.0. The checkpoint is served with SGLang and a MiMo chat template, and its SFT data totals 77.4B tokens.

  9. LMSYS OrgOfficialAI score65

    SGLang adds NVFP4 KV cache for longer context on Blackwell GPUs

    AILMSYS Org says NVFP4 KV cache in SGLang fits about 1.78x more context into GPU memory and speeds long-context decoding by up to 78%. Built with Alibaba Qwen and NVIDIA for Blackwell, it stores KV at about 56% of FP8's per-token footprint, with decode throughput up 37%, 58%, and 78% at 32K, 160K, and 1M context. The post reports near-lossless accuracy versus FP8 on GPQA-Diamond and AIME 2025 using Qwen3.5-397B-A17B, and it can be enabled with --kv-cache-dtype nvfp4.

    Why it matters: The post gives specific memory and throughput figures for NVFP4 KV cache in SGLang, showing how the format trades cache footprint against long-context decode speed.

    Image from @lmsysorg's post
  10. Xiaomi MiMo · new models on Hugging FaceOfficialAI score74

    Xiaomi MiMo-V2.6-Pro-RL released as 1.02T-parameter omnimodal model

    AIXiaomi MiMo released MiMo-V2.6-Pro-RL on Hugging Face, a sparse MoE model with 1.02T total and 42B activated parameters and a 1M-token context. The technical report says it accepts text, image, video, and audio, and was trained with a single mixed reinforcement learning run across coding, agent, visual, and cybersecurity tasks.

    Why it matters: The report pairs a 1.02T-parameter MoE model with an RL-based self-improvement method, useful for judging how reinforcement learning is scaled in frontier open models.

  11. ModelScopeOfficialAI score3

    ModelScope offers merch at Apsara 2026 booth in Hangzhou

    AIModelScope is promoting its booth at Apsara 2026, where visitors can get merchandise such as backpacks, tote bags, mugs, plush pendants, and hats. The booth is located at Booth 1-5C on the first floor of the Intelligence Engine hall at the Hangzhou International Expo Center Phase II.

    Image from @ModelScope2022's post
  12. Tim DettmersBlogAI score62

    Tim Dettmers argues academia can lead AI research with open-source local tools

    AITim Dettmers argues that agents make single research projects cheap, so the ecosystem, not the paper, becomes the unit of research. He says his lab's open-source week will release two projects and four papers together, including a harness that runs frontier-scale models on local hardware. He also argues that AI job fears are overstated and that university labs can compete on creativity and cheap, valuable problems.

  13. Interconnects (Nathan Lambert)BlogAI score65

    Chinese labs lead open-weight models in benchmarks, downloads, and research use

    AINathan Lambert argues that Chinese open-weight models now lead American ones on benchmarks, Hugging Face downloads, and OpenRouter usage. He estimates the gap to the American closed frontier at 2 to 5 months for Chinese open models and 6 to 9 months for American open models. The piece also reports that Chinese open-weight models were mentioned in over 40% of arXiv papers he scanned, compared with 30% for American models.

Sep 20

Sep 20Sun
  1. QwenOfficialAI score34

    Qwen-Image-2.1 now supported in ComfyUI for image generation

    AIQwen-Image-2.1 is now supported in ComfyUI, and Qwen invites users to try it and share their creations. ComfyUI describes it as an open-weights 7B checkpoint that handles both generation and editing, with native 2K image generation and instruction editing from up to 10 reference images in one pass.

  2. QwenOfficialAI score56

    Qwen-Image-2.1 releases open weights for image generation and editing

    AIAlibaba's Qwen team released Qwen-Image-2.1 as an open-weights image model for both generation and editing, with a lightweight 7B architecture. The model natively generates and edits RGBA layers, supports up to 10 reference images for editing, and is available on GitHub, ModelScope, and Hugging Face.

    Image from @Alibaba_Qwen's post
  3. ModelScopeOfficialAI score62

    Qwen-Image-2.1 unifies image generation and editing with native transparency

    AIAlibaba's ModelScope introduces Qwen-Image-2.1, a model that handles image generation and editing together, with native transparency and a compact 7B visual generation component. It adds KV cache reuse to speed up generation and editing while reducing memory use, especially with multiple reference images. The model can combine up to 10 reference images, make targeted local edits, and preserve portrait identity and product details.

    Why it matters: The post names concrete capabilities and a 7B size, letting readers compare it against the larger image models in the accompanying chart.

    Image from @ModelScope2022's post
  4. OpenBMBOfficialAI score35

    OpenBMB's Augury model improves on-device plant ID for farmers

    AIA developer's Augury plant identification model, built on an OpenBMB model, raised photo top-1 accuracy from 71.8% to 80.2% by merging duplicate species keys and adding PCA whitening. The next steps are reaching 90%+ accuracy and building a phone GUI so farmers can use it on-device.

  5. Qwen · new models on Hugging FaceOfficialAI score62

    Qwen releases open-source Qwen-Image-2.1 with a prompt rewriting model

    AIQwen has open-sourced Qwen-Image-2.1, a unified text-to-image generation and image editing model with a 7B-parameter visual generation component. The release also includes Qwen-Image-2.1-PE-T2I, a fine-tuned Qwen3.5-VL 9B model that rewrites brief image requests in any language into detailed English prompts with a recommended aspect ratio.

    Why it matters: The release pairs a 7B visual generation component with a separate prompt rewriting model, showing how a brief image request becomes a detailed English prompt before rendering.

Sep 19

Sep 19Sat
  1. StepFunOfficialAI score62

    StepFun Launches Step 5 Preview, a 600B MoE Model for Agentic Work

    AIStepFun has released Step 5 Preview, a flagship model for agentic work that it says delivers frontier-level performance in software engineering and professional knowledge work, with particular strength in finance. The model is a 600B total, 27B active mixture-of-experts design with a 1M context window and vision support. StepFun says it offers substantially lower task cost at comparable intelligence, and open weights are scheduled for October 15.

    Why it matters: The post pairs a cost-versus-intelligence chart with specs and a later open-weights date, so readers can judge the cost tradeoff against named competitor models.

    Image from @StepFun_ai's post

Sep 18

Sep 18Fri
  1. LM StudioOfficialAI score62

    LM Studio adds Qwen3.8-27B running at up to 144 tok/sec on M5 Max

    AILM Studio announced that Qwen3.8-27B runs at up to 144 tokens per second on an M5 Max MacBook Pro through its partnership with Inco Splash. The post claims up to 3× the decode speed of Ollama, 2× oMLX, and almost 4× when an agent fans out into sub-agents. Inco Splash is described as an open-source inference engine built for the model and Apple silicon, available through the linked LM Studio blog.

    Why it matters: The post reports a concrete decode speed on Apple silicon and compares it against named local inference tools, which helps readers judge local deployment performance.

  2. TinkerOfficialAI score31

    Jasper's guide shows how reward tweaks shape search agent behavior

    AIJasper Lu's new blog post walks through training a search agent with GRPO, showing how small reward function changes teach a model to avoid sloppy tool calls, prune unnecessary documents, and balance persistence against token efficiency. The post makes every rollout browsable and releases the code as open source, with the full process from learning rate sweeps to reward shaping documented.

  3. Google GemmaOfficialAI score22

    DiffusionGemma runs as a parallel decision model, faster than autoregressive generation

    AIGoogle Gemma's account says DiffusionGemma, running in a Jev-style decision setup, denoises an open canvas in one step rather than generating tokens sequentially, taking about 0.2 seconds on a DGX Spark. It says full bidirectional attention lets every option attend to the full context at once, and that the model inherits Gemma 4's spatial vision capabilities for visual and text decisions.

    Video from @googlegemma's post
  4. LMSYS OrgOfficialAI score16

    LMSYS releases SGLang SSD expert pack blog post

    AILMSYS Org published a blog post introducing an SGLang SSD expert pack, with the full details available on its website. The post itself gives no further technical specifics, so the summary is limited to the announcement.