Skip to contentSkip to stories

Updated

#Open-source ecosystem

Showing low-relevance items too. Hide low-relevance items

Sep 29

Sep 29Tue
  1. vLLMOfficialAI score23

    vLLM presents keynote and talks at PyTorchCon North America

    AIThe vLLM project announced a strong presence at PyTorchCon North America, with core maintainer and Inferact CEO Simon Mo giving the keynote on scaling open frontier inference infrastructure. Other vLLM maintainers, including Nick Hill and Red Hat AI engineers, will lead a developer session and talks on agentic inference and attention.

  2. OpenClaw🦞OfficialAI score70

    OpenClaw Enterprise launches as an open-source control plane for persistent agents

    AIThe OpenClaw Foundation announced OpenClaw Enterprise, an open-source enterprise control plane for persistent agents, in collaboration with Red Hat, NVIDIA, and OpenAI. The product is built to run on an organization's own infrastructure and will always be free for organizations to use.

    Why it matters: The announcement names its collaborators and deployment model, which helps organizations judge how the enterprise control plane would fit their own infrastructure.

    Image from @openclaw's post
  3. OpenBMBOfficialAI score72

    One-Shot OPD: One Training Query Matches Most of Full-Data Distillation Gains

    AIResearchers from Tsinghua NLP and collaborators show that on-policy distillation with a single training query recovers 87% of full-data gains on math, reaching 68.5 versus 69.8 by step 300. The paper attributes the slow progress to how fast the student absorbs the teacher's signal rather than to dataset size. Code and the paper are publicly available on GitHub and Hugging Face.

    Why it matters: The paper isolates training data from the algorithm, showing one query nearly matches full-data on-policy distillation, which reframes where post-training gains come from.

    Image from @OpenBMB's post
  4. ModelScopeOfficialAI score54

    IQuest-Q1 released as 320B MoE model for long-horizon coding agents

    AIModelScope announced IQuest-Q1, a 320B MoE model with 15B active parameters and a 512K context window for agentic coding. The post reports scores of 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 64.6 on DeepSWE v1.1, and 63.0 on NL2Repo, and says weights are released under the IQuest-Q1 License.

    Image from @ModelScope2022's post
  5. Rest of WorldNewsAI score46

    China's Open-Source AI Platforms Seek to Rival Hugging Face After Block

    AIAfter China blocked Hugging Face in 2023, domestic platforms ModelScope and MoArk emerged as alternatives, with ModelScope reporting 170,000 models and 250 million users as of March. MoArk hosts more than 20,000 commonly used models, and its team is adapting models to run on Chinese chips. Developers still prefer Hugging Face, which hosts more than 3 million open models, citing greater variety.

  6. ModelScopeOfficialAI score44

    Intern-Decision multimodal models scale structured decisions at 0.8B–4B

    AIShanghai AI Laboratory's Intern-Decision family of 0.8B, 2B, and 4B multimodal models averages 79.38, 84.68, and 90.02 across seven decision benchmarks. Intern-Decision-4B scores 88.74, surpassing Jev while achieving better probability calibration. Reported mean latency is 33.98, 33.28, and 44.16 ms, versus 109.70 ms for Jev in the same local HF setup.

    Image from @ModelScope2022's post
  7. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score40

    InternLM releases AdvancedMathBench-AutoVerifier to grade natural-language math proofs

    AIInternLM's AutoVerifier, built on Qwen3_5MoeForConditionalGeneration with about 68 GiB of weights across 40 safetensors shards, evaluates natural-language mathematical proofs, explains errors, and identifies the earliest incorrect step. It serves as the automatic grader for AdvancedMathBench's ProverBench, which accepts a proof only when all eight judgments report -1. The model is a learned grader rather than a formal proof checker and can make errors.

  8. OpenBMBOfficialAI score34

    MiniCPM-o 4.5 now runs in SGLang Omni v0.1.7 for developers

    AIOpenBMB announced that MiniCPM-o 4.5 is now supported in SGLang Omni v0.1.7, giving developers more flexibility to run and build with the model. The background release notes add that MiniCPM-o 4.5 brings multimodal input and speech output to the runtime. MiniCPM-o and MiniMax-Music3 also gained Intel XPU support in the same release.

  9. SGLangOfficialAI score53

    SGLang adds Day-0 support for IQuest-Q1 with a single-node serve command

    AISGLang says it has Day-0 support for IQuest-Q1, an open-source sparse MoE model with 320B total and 15B active parameters for coding and agentic tasks. The post includes a single-node serving command for H200 GPUs in BF16, using tensor parallelism of 8, EAGLE speculative decoding, and the iquest_q1 reasoning and tool-call parsers. The image marks the command as not verified.

    Image from @sgl_project's post
  10. Artificial Analysis ArticlesOfficialAI score62

    Artificial Analysis open-sources AA-AgentPerf-Local for benchmarking local AI agents

    AIArtificial Analysis has open-sourced AA-AgentPerf-Local, a tool that replays recorded agent trajectories to measure inference speed on laptops and workstations. Initial results cover NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with the RTX 5090 fastest for models that fit its 32 GB. The source states the tool and leaderboard will expand to more hardware, frameworks, and models.

    Why it matters: The source gives per-system completion times and memory bandwidth figures, letting readers compare local hardware for running agentic workloads.

Sep 28

Sep 28Mon
  1. ModelScopeOfficialAI score44

    Audio8 ASR Infinite enables unlimited-length streaming speech transcription with bounded memory

    AIAudio8 ASR Infinite transcribes Chinese and English audio of unlimited length using a rolling KV Cache that avoids accumulated drift. At a 480 ms delay, it reports 1.75 CER on AISHELL-1, 2.89 on AISHELL-4, and 3.04/6.81 WER on LibriSpeech test-clean/test-other. The preview release is under Apache 2.0, with deployment through an adapted vLLM stack.

    Video from @ModelScope2022's post
  2. vLLM BlogOfficialAI score54

    vLLM guide explains disaggregated serving for prefill and decode

    AIThe vLLM blog guide explains how separating prefill and decode, and moving tokenization to a CPU-only render tier, can keep token streams from stalling under load. In a two-L40S test on Qwen2.5-7B, collocated p99 inter-token latency reached 169 ms at 0.4 req/s while disaggregated serving stayed between 25 and 52 ms. The guide notes that the gain depends on fast KV cache transfer, and it includes setup code for NIXL-based serving and the render/derender API.

  3. World LabsOfficialAI score67

    Fei-Fei Li joins AMD as chief scientist as World Labs team joins

    AIFei-Fei Li will join AMD as Executive Vice President and Chief Scientist, working directly with CEO Lisa Su. World Labs will join AMD to form a frontier research organization, co-led by Justin Johnson and Ben Mildenhall, focused on an end-to-end open AI ecosystem spanning hardware, software, platforms, and widely accessible open models.

    Why it matters: The announcement shows how a leading AI lab's team is folding into a chipmaker, with a stated plan for an open AI ecosystem spanning hardware and models.

  4. Andrew NgXAI score46

    Andrew Ng says OpenWorker will use Nvidia OpenShell for sandboxed AI agents

    AIAndrew Ng says OpenWorker, his open-source agent harness for cybersecurity workflows, will run each agent's commands inside a sandbox built on Nvidia OpenShell. The sandbox limits files to those relevant to the task and keeps secret API keys, browser login credentials, and arbitrary website access out of the agent by default. Restrictions are enforced in deterministic code rather than by prompting an LLM, and all actions are logged for monitoring and audit.

  5. François CholletXAI score32

    K3-Node: a Keras 3 GNN library running on JAX, PyTorch, and TF

    AIK3-Node is a graph neural network library built natively on Keras 3, with models that run on JAX, PyTorch, and TensorFlow with hardware acceleration including Apple Silicon and TPU. According to the post, it achieves 100% public API parity with PyG and incorporates foundation models and architectures from Spektral and StellarGraph.

  6. Hugging FaceOfficialAI score14

    Hugging Face asks developers for Gradio feature requests

    AIHugging Face is inviting developers to suggest features for Gradio, its tool for building AI app frontends. The post follows a quoted message from Abid Lab noting that Gradio's original purpose has shifted and that the team is redesigning it from scratch around developers' current challenges with AI apps.

  7. RadixArkOfficialAI score46

    RadixArk releases Miles v0.1.1 with multi-LoRA and expanded model support

    AIRadixArk has released Miles v0.1.1, adding multi-LoRA with Tinker API compatibility so multiple training jobs can share one base model. The update also supports agentic training with harnesses like Claude Code and runs Harbor tasks in sandboxes including AgentENV, Daytona, E2B, and Modal. It further reduces memory needs for training larger models on validated NVIDIA and AMD GPUs and adds stable support for Qwen3.8-Flash-Next, GLM-5.3-Flash, and Kimi-K3.

    Image from @radixark's post
  8. Daniel HanXAI score29

    Unsloth Desktop serves local Laya decision models for real-time packing demo

    AIUnsloth Desktop can now serve local Laya decision models through a Jev-compatible API, shown in a real-time packing demo where suitcase items update as the user types. The demo runs through Unsloth's Decision API, and the team says more optimizations are coming to speed up local hardware performance.

    Video from @danielhanchen's post
  9. Google Cloud · AI & Machine LearningOfficialAI score40

    Why startups should pair open models like Gemma 4 with frontier APIs

    AIGoogle Cloud argues startups should combine open-weight models with frontier APIs rather than routing every request to one frontier model. It cites Gemma 4, which spans five sizes including a 31B dense model and a 26B A4B Mixture-of-Experts model, released under Apache 2.0. The article's examples report a 44% latency drop for Cue, from 876 ms to 488 ms, and a $0 server cost for BetterSpeak's on-device Gemma 4 E2B.

  10. clem 🤗XAI score49

    Hugging Face proposes egress usage monitoring for OpenShell agent sandboxes

    AIHugging Face is contributing egress usage monitoring to NVIDIA's OpenShell, part of the newly launched Open Agent Safety Platform, arguing that allowlists alone restrict where agents can go but not what they do. The proposed features include per-sandbox network budgets for requests, writes, and bytes, drift detection against each sandbox's baseline and cohort, and a fleet view that flags many sandboxes writing to one host even when every request is allowed.

    Video from @ClementDelangue's post
  11. Unsloth AIOfficialAI score34

    Laya Decision models can now run locally on 4GB RAM

    AIUnsloth AI says Laya Decision models can run locally on just 4GB of RAM, on CPU, Mac, Windows, Linux, and GPU setups. The post adds that Laya can be served through a Jev-compatible API via Unsloth Desktop.

    Image from @UnslothAI's post
  12. Lovable BlogOfficialAI score57

    Lovable apps can now run inside a company's Microsoft tenant

    AILovable announced a partnership with Microsoft that lets users publish apps into their company's Microsoft Entra tenant using Copilot Managed Runtime. Apps can connect to Microsoft 365, Fabric, Dataverse, and SQL data, and staff sign in with their work login. Copilot Managed Runtime is in public preview, and Microsoft 365 connectors, Fabric, and Microsoft sign-in are available on every Lovable plan, while Entra workspace sign-in is included on Business and Enterprise.

  13. ModelScopeOfficialAI score46

    Qwen-Image-2.1 LoRAs extract and remove layers for editing

    AIModelScope released two Qwen-Image-2.1 LoRAs, LayerExtract and LayerRemove, for layer-based image editing. LayerExtract isolates a prompt-specified subject onto a transparent background, while LayerRemove deletes the matching object from the source image and reconstructs the scene behind it. Both can be hot-swapped within the same DiffSynth-Studio pipeline, and the LoRA weights are licensed under Apache 2.0, with Qwen-Image-2.1 base-model terms also applying.

    Image from @ModelScope2022's post

Sep 26

Sep 26Sat
  1. clem 🤗XAI score31

    Xiaomi open-sources MiMo-V2.6 RL environments on Hugging Face

    AIXiaomi released an open-source RL environments repository, MiMo-V2.6-RL-oss, on Hugging Face. Hugging Face CEO Clément Delangue promoted the release, and a commenter estimated that comparable commercially sold tasks cost hundreds to thousands of dollars each.