Skip to contentSkip to stories

Updated

#Open-source ecosystem

Showing low-relevance items too. Hide low-relevance items

Sep 26

Sep 26Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score44

    Intern-Decision-2B: Structured Multi-Question Decision Model Fine-Tuned from Qwen3.5-2B

    AIShanghai AI Lab's InternLM released Intern-Decision-2B, a multimodal structured decision model fine-tuned from Qwen3.5-2B that returns calibrated answer distributions for multiple questions in one forward pass. It averages 84.68 across listed benchmarks with a 0.437 Brier score and 33.28 ms mean latency on a single RTX 4090. Model weights, a Python DecisionEngine API, and GitHub code are available, with support for up to 16 questions and eight images.

  2. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score46

    Intern-Decision-0.8B: InternLM's structured decision model on Hugging Face

    AIInternLM released Intern-Decision-0.8B, a multimodal structured decision model fine-tuned from Qwen3.5-0.8B that scores answers to multiple questions in one forward pass. The model reports a 79.38 average score and a 33.98 ms mean latency on a single RTX 4090, with 0.8B, 2B, and 4B sizes available. It is accessed through a Python DecisionEngine API that returns calibrated probabilities rather than generating free-form text.

Sep 25

Sep 25Fri
  1. LMSYS OrgOfficialAI score38

    SGLang adds multi-item scoring for faster decision model serving

    AISGLang's /v1/score endpoint returns scores for exact requested labels such as Yes/No or A/B/C, and its multi-item scoring (MIS) computes shared context once while keeping candidates isolated. On Qwen3-8B, 16-candidate p95 latency dropped from 54.1 ms with Generate to 20.6 ms with MIS. On Qwen3-0.6B, MIS p95 stayed under about 100 ms as load rose, versus seconds for Generate and SIS.

    Image from @lmsysorg's post
  2. GitHub Blog · AI & MLOfficialAI score33

    How to build custom workflows with canvases in the GitHub Copilot app

    AICanvases in the GitHub Copilot app are customizable interfaces that you and the agent share, such as kanban boards, dashboards, or checklists. You create one by running /create-canvas and describing the workflow, what you can do in the interface, and what the agent can do. Changes made by either you or the agent appear immediately in the shared canvas, and completed canvases can be saved as reusable extensions.

  3. Together AIOfficialAI score12

    Together AI Simplifies Access to Frontier Open Models via API

    AITogether AI says teams can access frontier open models through its API without managing underlying infrastructure, a point Ted Cui, its VP of Engineering and Inference Platform, made at Apsara Conference 2026. The company emphasizes reliable, fast inference as essential for this access.

Sep 24

Sep 24Thu
  1. vLLMOfficialAI score30

    vLLM integrates TileRT with PD disaggregation, benchmark config published

    AIvLLM has published a blog post explaining how its integration with TileRT works, alongside a public benchmark configuration in the InferenceX repository. The benchmark script covers GLM-5.3 FP8 on MI355X hardware and is linked on GitHub. The post itself provides the integration details.

  2. vLLMOfficialAI score42

    TileRT and vLLM hit 469 tok/s on GLM-5.3 with MI355X

    AIThe TileRT and AMD teams reached 469 tok/s single-user decode for GLM-5.3 on 8× MI355X using vLLM. The setup disaggregates work, with vLLM handling prefill and TileRT handling latency-critical decode through vLLM's V1 connector interface. SemiAnalysis's AgentX benchmark reports the configuration at 470 TPS on GLM 5.3 (FP8), over 40% faster than GB300 TRTLLM using FP4.

  3. Lewis Tunstall @ COLM 🌉XAI score42

    Hugging Face releases over 5,000 RL environments for data science tasks

    AIHugging Face released SmolDataEnvs, more than 5,000 open-source RL environments aimed at real-world data science tasks. They target the gap between simple educational games and frontier-level benchmarks, especially for improving coding in models under 10B parameters. The environments are designed as a testbed for developing new RL methods such as GRPO or OPSD.

  4. GitHub Blog · AI & MLOfficialAI score66

    GitHub Security Lab shows an LLM agent running AI-driven fuzzing for C/C++ projects

    AIGitHub Security Lab describes the Fuzzing Taskflow, an LLM agent pipeline that identifies entrypoints, writes harnesses, runs AFL++, reads coverage reports, and triages crashes for C/C++ repositories. The agent makes decisions while MCP tools handle execution, and state is stored in a SQLite database. The post also warns that the taskflow runs AFL and build commands directly on the host, so it should be used only in disposable environments without elevated privileges.

    Why it matters: The post explains how an LLM agent automates fuzzing steps like harness writing, coverage gap chasing, and crash triage, with a runnable workflow and design tradeoffs.

  5. Google for DevelopersOfficialAI score37

    Gemma 4 now runs on-device in the Antigravity SDK

    AIGoogle says Gemma 4 can now run locally on-device within the Antigravity SDK. Developers can build fully local or hybrid multi-agent workflows that pair cloud models with Gemma 4 agents for auditing, patching, and testing code. The post emphasizes total data privacy and zero API fees, powered by LiteRT.

    Video from @googledevs's post
  6. OdysseyOfficialAI score18

    Odyssey introduces Agora-2, a multi-agent world model

    AIOdyssey launched Agora-2, a multi-agent world model that the company is making available for public experimentation. The post predicts such models will increasingly power applications in AI training, AI safety, robotics, autonomous vehicles, defense, energy, cybersecurity, and gaming.

  7. ModelScopeOfficialAI score38

    Qwen-Image-2.1-Fun-Controlnet-Union adds eight controls and inpainting

    AIModelScope released Qwen-Image-2.1-Fun-Controlnet-Union, a single checkpoint adding eight structural controls, including Canny, Depth, Pose, and Scribble, plus inpainting to Qwen-Image 2.1. Control and inpainting share one branch with 16 injection points across every second Transformer block, keeping the base model frozen and requiring no checkpoint switching. It runs at guidance scale 1.0 with CFG-distilled sampling and prefix KV caching, and is available under the Qwen Research License with base Qwen-Image 2.1 weights required.

    Image from @ModelScope2022's post
  8. Goodfire ResearchOfficialAI score52

    Block-Sparse Featurizers Recover Multidimensional Concept Geometry in Vision Models

    AIGoodfire Research introduces Block-Sparse Featurizers (BSF), which decompose model activations into subspaces rather than single directions. Applied to DINOv3 and Stable Diffusion XL, BSFs find interpretable multidimensional features that better explain activations and enable fine-grained steering. The authors report that most concepts they examined have a stable rank of about two to four dimensions.

Sep 23

Sep 23Wed
  1. Liquid AI BlogOfficialAI score46

    LFM2.5-VL-DSpark speeds up vision-language model decoding on GPUs and edge devices

    AILiquid AI released an experimental DSpark draft model for its LFM2.5-VL-3B vision-language model, delivering decoding throughput gains of up to 2.66× on GPUs and 3.13× on edge devices. The drafter adds about 280M parameters, an 8.9% increase in the deployed model's parameter count, and is available on Hugging Face with support in llama.cpp, SGLang, and MLX-VLM.

  2. vLLM BlogOfficialAI score54

    vLLM adds distortion-free Gumbel-max watermarking for text provenance

    AIvLLM now supports Gumbel-max watermarking, which embeds a keyed signal into generated text without changing the expected token distribution. Detection requires the secret key and tokenizer, and the signal accumulates over longer outputs. Benchmarks on Qwen3.5-27B with MTP-3 show throughput changes between -1.1% and +2.0% across batch sizes, with no consistent slowdown.

  3. Google Developers BlogOfficialAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    AIGoogle Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    Why it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  4. Google GemmaOfficialAI score60

    Google's Antigravity SDK adds local execution with Gemma 4 and LiteRT

    AIGoogle says the Antigravity SDK now supports running agents entirely on a local machine with Gemma 4 and LiteRT. The post adds support for OpenAI-compatible endpoints, naming Ollama, llama.cpp, and vLLM as options for serving Gemma, and gives the install command pip install google-antigravity litert-lm.

    Video from @googlegemma's post
  5. InferactOfficialAI score44

    vLLM maintainers show TPUv7 megakernels beat GB200 NVL72 on Kimi K3

    AIInferact says vLLM maintainers used megakernel optimization to reach 700 tokens per second per user on TPUv7 running Kimi K3. SemiAnalysis, which shared the work, reports this is 56% better performance than Nvidia's GB200 NVL72. Inferact links a full technical breakdown of the TPU megakernel work on its blog.

  6. Google AntigravityOfficialAI score38

    Antigravity SDK runs Gemma 4 fully offline on local GPUs

    AIGoogle Antigravity says developers can now run open models such as Gemma 4 completely offline in its SDK. The setup uses Google AI Edge's LiteRT to run the model directly on a local GPU, with no API costs and no internet connection required.

    Video from @antigravity's post
  7. InferactOfficialAI score49

    Inferact's TPU megakernel runs Kimi K3 at 709 tokens/s

    AIInferact says its first TPU megakernel for Kimi K3 reaches 709 tokens/s on low-concurrency decode with DSpark speculative decoding, versus 450 tokens/s for its GB200 baseline. The company claims it is the first TPU inference megakernel, running the whole model in a single Pallas kernel, and says it is roughly 1.4 to 2x the GB200 baseline at batch sizes 1 through 8 without speculative decoding. Inferact says it is open-sourcing the kernel today.

    Video from @inferact's post
  8. LM StudioOfficialAI score28

    Bionic adds a built-in interactive canvas for shared diagrams

    AIBionic now includes a built-in interactive canvas where users can create Excalidraw diagrams that both they and Bionic can view and edit. The canvas supports collaboration on mockups, system designs, and process maps, and users can ask Bionic to implement what is drawn.

    Video from @lmstudio's post
  9. ModelScopeOfficialAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

    Image from @ModelScope2022's post
  10. ModelScopeOfficialAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    AIShanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

    Image from @ModelScope2022's post

Sep 22

Sep 22Tue
  1. Daniel HanXAI score22

    Unsloth Desktop hotfix adds Qwen-Image-2.1 image editing and fixes

    AIUnsloth Desktop received a hotfix update adding image editing for Qwen-Image-2.1. The update also fixes diffusers update issues, GGUF loading failures for Qwen-Image, and black artifacts during diffusion on A100 and consumer GPUs. Users should receive a banner prompting them to update.

  2. Fireworks AI BlogOfficialAI score46

    Fireworks ARCv3 cuts RL weight-update payloads nearly 50% for cross-region training

    AIFireworks released ARCv3, a lossless compressor for BF16 weight-update deltas sent from trainers to RL rollout machines. Across 1,000 production RL deltas, ARCv3 produced payloads nearly 50% smaller than ARCv2, averaging about 0.19% of the BF16 weight size versus 0.36%. ARCv3 is available through the Fireworks Training API as fireworks-delta-compression.

  3. Boris ChernyXAI score42

    Boris Cherny Uses Opus 5.5 to Formally Verify Claude Agent SDK

    AIBoris Cherny used Opus 5.5 to formally verify the Claude Agent SDK with Lean, and short prompts produced 16 PRs fixing bugs and race conditions. He also combines Lean and TLA+ to find issues in data flow, concurrency, and state management, and says Claude is strong in both languages even though he does not know them well.

    Video from @bcherny's post
  4. whXAI score34

    MiMo-V2.6 paper details data and RL results for open model

    AIThe MiMo-V2.6 paper thread reports on the newest open model, which also streams its RL run, focusing on data and RL experimental results rather than architecture. The Pro model reportedly rose from 58.41 to 72.57 on DeepSWE after RL, with the top published DeepSWE score cited at 74.

    Image from @nrehiew_'s post
  5. Comfy BlogOfficialAI score42

    ComfyUI Speeds Up MiniMax H3 Video VAE Encoding and Decoding

    AIComfyUI's update makes the MiniMax H3 video VAE encode up to about 2.2x faster and decode 1.4-2.7x faster, cutting a 1344x768, 129-frame round trip on an RTX 5090 from 24.3 to 12.7 seconds. The gains come from a fused encoder kernel enabled by default, fp16 accumulation support in a custom convolution, and an int8 decoder, and the source says the changes are visually lossless to the eye. Users need ComfyUI v0.36.0 or above, and the int8 VAE file is a drop-in replacement for the standard one.

  6. Daniel HanXAI score42

    Qwen-Image-2.1 runs locally in Unsloth Desktop via INT8, FP8, GGUF

    AIDaniel Han says Qwen-Image-2.1 works in Unsloth Desktop through INT8, FP8, and GGUF builds, with Unsloth also releasing dynamic GGUFs for it. Pinned RAM offloading lets INT8 and FP8 fit under 6–8GB of VRAM while remaining relatively fast. The linked Unsloth post says the 7B model runs on 12GB VRAM and performs on par with Nano Banana 2.0.

  7. Unsloth AIOfficialAI score70

    Qwen-Image-2.1 runs locally on 12GB VRAM using Unsloth GGUFs

    AIUnsloth says the 7B Qwen-Image-2.1 text-to-image and editing model can run locally on 12GB VRAM using its GGUF builds. It also states that the model performs on par with Nano Banana 2.0, and that Dynamic FP8 can run on 6GB of VRAM via offloading for higher quality. The image lists int8 at 7.26 GB with mean LPIPS 0.064 and fp8 at 7.12 GB with mean LPIPS 0.112, and says int8 is the default.

    Why it matters: The post gives concrete local-run settings, VRAM figures, and GGUF and FP8 options, which helps readers judge whether the model fits their hardware.

    Image from @UnslothAI's post
  8. Sebastian RaschkaXAI score62

    Xiaomi MiMo-V2.6-Pro tops open-weight benchmarks with simple attention design

    AIXiaomi's MiMo-V2.6-Pro ranks first among open-weight models on the Artificial Analysis Intelligence Index with a score of 46. The author attributes its standing mainly to a training data and post-training recipe that increased agent tasks and used an agentic grader for rewards, rather than its plain Grouped Query Attention and Sliding Window Attention design with a 128-token window.

    Image from @rasbt's post