Updated
#Open-source ecosystem
Updated
Oct 8
Nous ResearchAI score14 OpenBMBAI score36 ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%
AIReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.
Clément DelangueAI score22 We urgently need more public traces of AI agents attacking and defending systems.
AIDefenders can’t learn from what they can’t see. If you have traces and are being pressured to keep them private, my DMs are open. Let’s level the playing field and fight the asymmetry and lack of transparency in AI!
StepFunAI score62 StepFun releases Step 5 Preview on OpenRouter with a week of free access
AIStepFun's Step 5 Preview is now live on OpenRouter, with a week of free access rolling out across opencode, Cline, NousResearch, and Kilo Code. The company positions it as flagship-tier intelligence for agentic and professional work at substantially lower task cost, letting users keep their workflow while switching models.
Nathan LambertAI score18 The South Korean bank hack fits with this.
AIYes, open models are being fitted to be used by some attackers, but so long as Claude Code can be the orchestrator of a hack we desperately need open models to diffuse defense as well. A hard balance societally to stomach, but needed.
PyTorch BlogAI score46 IBM Builds Spyre as a Native PyTorch Device via torch-spyre
AIIBM's torch-spyre integration makes Spyre, its dataflow inference accelerator, a native PyTorch device by mapping PyTorch's device, allocator, stream, and event abstractions onto the Spyre runtime and firmware. Tensors stay resident on device="spyre" between operations, and FX graphs remain in the Inductor compiler path. The approach gives eager and compiled execution one path with lower launch overhead.
JetBrains AI BlogPickAI score62 JetBrains releases Mellum2.1, an open coding model trained with reinforcement learning
AIJetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.
Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.
meng shaoAI score55 Tencent Cloud open-sources Octop, a self-hosted multi-agent AI assistant platform
AITencent Cloud has open-sourced Octop, a self-hosted multi-agent AI assistant platform aimed at families and small teams, with multi-user accounts and data kept on the user's own machine. The full text describes it as a single Python process that bundles the backend, web dashboard, CLI, IM gateway, cron jobs, and multi-agent runtime, with state rebuilt from SQLite on restart.
QbitAIAI score44 PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End
AIUniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.
Anthropic NewsroomPickAI score62 Anthropic launches Cyber Mission with infrastructure defense and free OSS Scanner
AIAnthropic has launched the Anthropic Cyber Mission, which starts with the Critical Infrastructure Defense Program for operational technology and OSS Scanner for open-source projects. The defense program brings frontier Claude models, on-site engineers and threat research to trusted providers such as Accenture, CrowdStrike and Palo Alto Networks. OSS Scanner gives enrolled open-source projects periodic free scans from its strongest models, with reports sent without human review and an expected true-positive rate above 90%.
Why it matters: The announcement shows how a frontier AI lab is packaging cyber defense around critical infrastructure and open-source maintainers, including the program's partners and access routes.
Anthropic ResearchAI score72 Anthropic launches OSS Scanner, a free AI vulnerability scanner for open-source projects
AIAnthropic is launching OSS Scanner, an opt-in service that runs periodic security scans of enrolled open-source projects using its strongest models at no cost. Its outputs are fully model-generated without human review, so some reports may be incorrect or invalid, though a pilot found 85 of 97 checked critical and high-severity findings met Anthropic's disclosure bar. Core maintainers of eligible projects can enroll through a GitHub pull request.
Oct 7
meng shaoAI score75 Microsoft positions Windows as the home for hybrid AI agents across four layers
AIMicrosoft has repositioned Windows as the home for hybrid intelligence, where AI agents can run locally or in the cloud. The announcement covers four layers: MXC reaching general availability for agent isolation, local models such as MAI Code 1.1 Flash, Copilot on Copilot+ PCs gaining local context and actions in coming months, and new hardware including RTX Spark PCs and DGX Station for Windows.
Orange AIAI score34 Next Token episode 5 covers Personal Agents, open-source software, and hardware projects
AIThis Next Token episode discusses Personal Agents, including Dots in Codex, memory and cloud computer permissions, and whether agents should act as assistants or digital twins. The hosts also cover Instinct's booking and business-travel model, hands-on projects built with Opus 5.5, and whether software, games, and hardware could become open source as AI makes rewriting easier.
MarkTechPostAI score58 Unsloth Studio re-checks changed model repos and blocks flagged weights before loading
AIUnsloth Studio binds remote-code approval to a fingerprint of the scanned code, so changed code requires fresh consent before it runs. It also blocks weight files that Hugging Face has flagged for malware in the path the selected loader would deserialize. The article describes these checks as one layer among several, alongside package-content scans and OS sandboxes, and notes that the scanner is not a sandbox and cannot catch every evasion.
AI EraAI score54 Mistral Large 4 launches with 1 trillion parameters and planned open weights
AIFrench AI company Mistral has released Mistral Large 4, its new flagship model with 1 trillion total parameters. The company plans to make all of the model's weights open at the end of the month.
Epoch AIPickAI score67 Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it
AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.
Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.
Hugging Face BlogPickAI score66 How one developer built six custom models with ML-Intern for about USD 103
AIA Hugging Face blog author used the ML-Intern agent in HuggingChat to build six small models by writing detailed prompts that specify datasets, base models, baselines, smoke tests, and spending limits. The projects include a citrus disease vision-language model, a Huggy character LoRA, a camera-angle LoRA, a doodle-to-object LoRA, a 0.8B prompt rewriter, and a 4-step distilled Agate model, with total compute cost of about USD 103. Each project's prompts and public models are linked from the post.
Why it matters: The author shows how prompt structure, baselines, smoke tests, and budget caps shape an agent-driven training workflow, with per-project costs given.
TekniumAI score28 hello
Ars Technica · AIAI score46 Artcraft releases open source clones of Adobe Photoshop, Premiere and other apps built with Claude
AIDeveloper Brandon Thomas's Artcraft has launched seven open source apps in Rust that aim to replicate the interfaces and tools of Adobe Photoshop, Illustrator, Premiere, Lightroom, After Effects, InDesign, and Acrobat Pro. Thomas said he used Anthropic's Claude Opus 5.5 to build the clean-room replacements, with WebAssembly versions available for browser use. The apps remain in a "super early alpha" state, and commenters have pointed out many current shortcomings.
Together AIAI score14 "It's well worth spending the money on tokens if the business is getting the benefit." — IBM Cloud GM Alan Peacock We're scaling…
AI…open-source inference with @IBM + @nvidia on a dedicated B300 cluster on IBM Cloud. Alan + our CRO Kai Mak 👇🏻
OpenRouterAI score46 Perplexity Decider v1.1 is now on OpenRouter @perplexity_ai's open-weights multimodal decision model takes text, JSON, or images and…
AI…returns typed answers with probabilities $0.02/M input. Output is free
vLLMAI score22 4/ Together, TTFT drops nearly 70% at ~100K throughput.
AIThanks to @deepseek_ai for the model and kernels, @nvidia for the collaboration, and @SemiAnalysis_ for AgentX. Built by @inferact and the vLLM community.
vLLMAI score46 1/ Three weeks from day 0, DeepSeek-V4.1-Flash on vLLM runs 1.9× faster at low concurrency and delivers 5.3× the throughput at 150 TPS per…
AI…user on @SemiAnalysis_ AgentX. Here is how, with interactive figures you can step through 🧵
Clément DelangueAI score34 If you need any more evidence about how exciting local AI (and llamacpp) is!
AIFrom the @Microsoft windows launch today with @satyanadella @sriramk @JensenHuang
ClaudeDevsAI score22 Use a compatible driver from @browser_use, @browserbase, @e2b, @daytonaio, or write your own based on the example drivers in the…
AI…quickstarts. Quickstart: Docs:
GitHubAI score57 GitHub Copilot local sandboxing becomes generally available
AILocal sandboxing for GitHub Copilot is now generally available. It lets Copilot run commands in an isolated environment with controlled access to files, networks, system capabilities, and credentials. Enterprise teams can also centrally manage policies, and the feature is available in GitHub Copilot CLI, the GitHub Copilot app, and @code.
Ai2AI score12 What’s happening at #COLM2026?
AIAi2 Comms Lead @Kyle_L_Wiggers catches up with Senior Director of NLP @nlpnoah to talk agentic models, the next iteration of Olmo, & more! At COLM? Come say hi at booth 302! 👋
OpenAI DevelopersAI score6 A one-shotted hockey game
Georgi GerganovAI score31 It is very cool to see llama.cpp on the big stage in todays Windows event!
AIThe software and hardware stacks are finally coming together. Our community has put a lot of hard work in the past years and it shows. Looking forward to more users embracing local AI.
MarkTechPostAI score60 Liquid AI releases open-weight d1-3B and d1-omni-600M decision models
AILiquid AI released two open-weight multimodal decision models, d1-3B and d1-omni-600M, which return probability answers in one forward pass with zero output tokens. d1-3B scores 48.57 on Decision Index v0.2.1 and answers one question in 8 ms on an RTX 4090, while the models are licensed free for commercial use below $10 million in annual revenue.
OpenRouterAI score38 Cloudflare Clef decision models are available on OpenRouter @Cloudflare's Clef (27B) and Clef Flash (9B) are open-source decision models.
AISend text, JSON, or images and get back typed answers with probabilities instead of generated text $0.24/M input for Clef, $0.09/M for Clef Flash. Output is free
Georgi GerganovAI score44 llama.cpp can distribute inference on heterogeneous devices through the ggml RPC backend It's an advanced setting but I think with time…
AI…we'll make it more accessible to regular users.
Liquid AIAI score38 Both Open d1 models run across NVIDIA DGX, RTX, and Jetson, with day-one llama.cpp support to run anywhere.
AId1-3B single-question latency, measured one request at a time: > NVIDIA RTX 4090: 8 ms > Jetson AGX Thor: 16 ms > Jetson AGX Orin 64 GB: 26 ms > Jetson Orin Nano: 50 ms 4/
Liquid AIAI score23 d1-3B ranks first among models under 10B on the Decision Index v0.2.1, a benchmark for structured decision-making.
AIBuilt from LFM2.5-VL-3B, it makes decisions from text and images in one pass. Use it for reranking, agent guardrails, and visual inspection. 2/
Nous ResearchAI score38 As reported in the @WSJ, we have raised a Series B to bring Hermes Agent to new frontiers (and, yes, build a mobile app) Thank you to our…
AI…investors including @nvidia @M12vc @SamsungNext @robotventures @usv @ycombinator @MenloVentures, among others
Daniel HanAI score48 We made it possible to train your own Decision model locally on just 3GB VRAM!
AIWe converted Qwen, Gemma, Llama all into decision models by fine-tuning using a Clef head, boosting accuracy from 30 to up to 78%. You can try it yourself via Unsloth Desktop!
LlamaIndexAI score47 LlamaIndex launches OpenDocRouter, one API for many document parsing models
AILlamaIndex announced OpenDocRouter, a single API that routes document parsing requests to any of 10 frontier and open-source models at launch, including Claude Opus 5.5, Gemini 3.8 Flash, GPT-6 Luna, MinerU2.5-Pro, and PaddleOCR-VL-1.6. Users can switch models in one line with the same request and markdown output, and each model is scored on ParseBench for quality and cost. Pricing is per-token, failed pages are not charged, and the service costs $0.86 to $48.82 per 1,000 pages depending on the model.
Ai2AI score36 We're making new Stage 1 checkpoints available with the byte-level components already trained & the original model's weights unchanged.
AIResearchers can build on these to test new architectures & train the full system without repeating that initial stage.