Simon Willison has released ttok 0.4, a command-line tool for counting tokens built on OpenAI's open source tiktoken library. The update adds a --list-models command, fixes a Click warning, and updates CI, and the tool can run through uvx, for example with cat file.txt | uvx ttok.
Theo, creator of the T3 Stack, open-sourced tsc-rs, a line-by-line Rust port of Microsoft's Go-native TypeScript 7 compiler, type checker, and language server under MIT, pinned to typescript-go commit 673a5f17. The author reports tsc-rs is about 1.61× faster than tsc 7 and about 2.95× faster than bun check on six real-app benchmarks on an Apple M4 Pro. The port passes all 181,711 ported Go tests, and CLI output matches the Go version on 120 open-source repos except for known edge cases such as monorepo rootDir and tsc -b incremental output.
Why it matters: The post reports a benchmarked, test-verified Rust port of the TypeScript 7 compiler, with pinned upstream and stated edge cases useful for judging its compatibility.
Shanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.
Laya is a 421-million-parameter non-autoregressive decision engine from Convai Innovations that returns calibrated option probabilities in a single forward pass with zero output tokens. This tutorial tests its zero-shot accuracy, probability calibration, temperature fitting, and abstention gating on the CLINC150 banking intent dataset.
Xiaomi MiMo has released MiMo-V2.5-ASR, an open-source speech recognition model that the company says achieves state-of-the-art results across multiple benchmarks. The model supports bilingual Chinese–English recognition, Chinese dialects such as Wu, Cantonese, Hokkien, and Sichuanese, code-switching, and lyrics transcription. It is also designed to handle noisy environments and multi-speaker conversations.
Xiaomi introduces the MiMo-V2.6 series of models, as stated on its official MiMo page. The source excerpt gives only the tagline, describing frontier intelligence across all modalities, so specific capabilities, benchmarks, and availability cannot be confirmed from it.
Atomic Agent Desktop is a free open-source app for macOS, Windows, and Linux that runs open models like Qwen and Gemma locally without an account. It connects to a cloud model only when selected, and its Fusion feature lets a cloud model plan a task while up to 8 local agents carry it out. The post's own text adds a setup wizard that checks RAM and suggests suitable models, and import from Claude Code, Codex, Hermes, and OpenClaw.
Databricks has released Funke, a Python and PySpark library and deployable pipeline that parses HL7v2 healthcare messages into native Spark types while preserving the full message hierarchy. It succeeds Smolder, the Scala data source Databricks open-sourced in 2021, and ingests through Auto Loader into Unity Catalog bronze and silver tables. Users can query segments, fields, components, and subcomponents directly with DataFrame or Spark SQL expressions.
whoops I hit send and went to a meeting, didn't realize the repo wasnt public- it is now! it's a chrome extension I use every day now https://github.com/ThariqS/ai-newtab
Anthropic is launching OSS Scanner, a service that uses its frontier models to periodically scan opted-in open-source projects for vulnerabilities at no cost. Its reports provide a proof-of-concept, an explanation, and a suggested fix.
it's built on top of Claude Managed Agents, doing code generation in the cloud sandbo, using Sonnet 5.5 each run costs ~$0.75 you can try it yourself here: https://github.com/ThariqS/ai-newtab/
Claude Code v2.1.295 adds onFailure: "block" for command and HTTP hooks, so a hook that cannot start, times out, or exits unexpectedly blocks the action. The release also adds an optional models list for Claude apps gateway upstreams, plus upstream_request_id in the inference audit event, and fixes a range of MCP, plugin, and terminal issues.
Claude Dashboards lets users connect a data platform or CRM tool and ask questions in plain language. Claude writes the query and builds a dashboard that updates as the data changes, and every chart shows its underlying query. The feature is in beta on paid plans.
OpenAI's Codex 0.162.0 release adds tools for creating and listing managed Git worktrees from trusted local projects when the worktrees feature is enabled. The update also lets users pin tasks in the agent Command Center, copy transcript blocks with /copy, and make URLs clickable in approval headers, questions, and warnings, along with several Linux and Windows sandbox fixes.
Mozilla.ai's cq project proposes a shared knowledge layer where AI agents capture lessons from non-obvious fixes as structured knowledge units that other agents can later query. The default setup is local-first, using a local SQLite database so nothing leaves the machine, with an option to connect to a remote team server that adds review.
Among models with a Hallucination-Gated All-Pass Rate above 0%, four set the Pareto frontier for score vs. Cost per Task: GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max) and Grok 4.7 (xhigh). Grok 4.7 (xhigh) leads at ~$9.50 per task and Muse Spark 1.3 (max) comes second at ~$4.20, while the three Claude models cost ~$18 to ~$22 per task. GPT-6 Luna (max) is the cheapest at ~$0.22 per task, scoring 3.3%.
Databricks introduced Vibe Data Modeling, an open-source agent that helps teams build, validate, and evolve business-specific data models. It applies roughly 250 modeling rules while keeping data modelers and business stakeholders involved. Teams can start from 40 industry models as a baseline and iterate toward models that reflect how their business operates.
if you use Grok Bot on Omarchy, or are building plugins for it, please let me know if you have any feedback or feature requests! Would be cool to see what interesting integrations we could support https://plugins.omarchy.org/?q=grok+bot#catalog
NVIDIA Dynamo uses a unified session-level identifier to make its inference stack aware of agent sessions, subagents, and their KV cache across turns and tool calls. On SWE-bench, two TP4 MiniMax-M2 replicas on one 8xH100 node gained roughly 12-16% throughput from program-aware scheduling over KV-aware routing alone. The post also describes experimental shared-pool indexing and a proposed KvHint interface for session-aware cache policies in vLLM and SGLang.
Why it matters: The post explains how session identifiers let an inference stack track agent working sets, with measured throughput gains on SWE-bench and agentic RL rollouts.
RSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.
We added OS level sandoxing in Unsloth with bwrap (Linux), seatbelt (Mac) and Windows MXC in Unsloth! Latency per tool call for all is under 100ms. Our software style sandboxing with regex ast checks is 3ms latency as well. Thanks to Windows for collabing with us on MXC!
Albums are live on Suno 🎵 Bring your songs together into a full release, set the artwork, arrange the tracklist, and publish when you’re ready. Already using a playlist as an Album? Turn it into one without rebuilding everything from scratch. #Suno #SunoAlbums
OpenClaw v2026.9.9 patch release is out 🦞 🧠 GPT-6.1 Sol in Codex + Claude Haiku 5.5 🔧 Better failed-update recovery 💬 Missing iMessage replies fixed ⏰ Scheduled-job fixes Thanks to all 90 contributors! https://docs.openclaw.ai/releases/2026.9.9
JetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.
Windows now has sandboxing! Microsoft released an open-source repo, mxc, for sandboxed code execution. We collaborated with Windows to add mxc OS level sandboxing to Unsloth which adds just <100 ms of overhead. GitHub: https://github.com/unslothai/unsloth Guide: https://unsloth.ai/docs/new/studio/sandboxing-in-unsloth
Carbon-A is an open model that predicts gene locations directly from DNA, and it has been used to annotate genomes from over 22,000 species. The release includes a database of 566 million gene candidates, about 16 times the gene annotations in the RefSeq dataset. Wet-lab RNA experiments supported 239 candidates missing from RefSeq across cats, Syrian hamsters, chickens, and Arabidopsis.
Why it matters: The source ties an open gene-annotation model to specific wet-lab checks and gene counts, helping readers judge how far its predictions extend beyond well-studied genomes.
Don't Worry About the Vase (Zvi Mowshowitz)AI score6262
OpenAI reportedly posted solutions to 90 of the top 500 open math problems, using an average of three hours of Pro-level compute per question. Anthropic released Claude Haiku 5.5 at $0.10 input and $0.50 output per million tokens, and the author says Jay Clayton was named AI Czar to head a new taskforce.
The team released Carbon-A, an open model that finds genes directly in DNA, along with a database of 566.34 million candidate genes across 22,617 species. The model reads genomes without needing a close relative, and wet-lab validation in cats, chickens, and arabidopsis is cited, with 239 genes found missing from reference annotations of common species. The authors say the model marks gene locations but does not design DNA or predict gene function.
New open source cross-platform (Windows, macOS, Linux) sandboxing library from Microsoft - looks very promising, uses processcontainer/bubblewrap/seatbelt under the hood https://github.com/microsoft/mxc
SynthID Detector (http://synthid.com) is now publicly available, supporting the Nano Banana 2.1, @OpenAI, @nvidia, Kakao, and soon @Apple. 🌐 Verify if an image, video, or audio file was generated: 1️⃣ Upload or paste your file at https://synthid.com 2️⃣ Scans for watermarks from Google or our partners (files are deleted right after)
The author says Opus 5.5 keeps improving and works well for producing weekly product short videos, with all materials generated directly without extra services. GoodCase added 269 new AI showcase cases, prompts, and 7 new Skills, bringing its total to 1,699 cases, 95 Skills, and 426 creators. The post also highlights awesome-seedance, which now lists 795 video cases, 367 prompt retests, 27 prompt templates, and 77 installable video Skills.
Your next build is on us. Step 5 Preview is now free in @opencode for one week. Our flagship model for agentic and professional work. Keep your workflow. Switch your model.
Amazon Web Services launched an open-source Physical AI Toolchain that combines AWS services with NVIDIA's Physical AI software to cover data generation, model training, simulation, edge deployment, and continuous improvement for robots. AWS uses Amazon SageMaker for training and AWS IoT Greengrass for distributing models to edge devices, while NVIDIA contributes Isaac Sim, Isaac Lab, Isaac GR00T, and Cosmos. The toolchain is hardware-neutral and does not directly replace RoboMaker, which was shut down in 2025.
new Llama.cpp release ships with (multimodal!) Jev-like models support, performance upgrade for Metal and more! 🔥 super simple: llama serve -hf ggml-org/Clef-Flash-GGUF browse all the decision models here https://huggingface.co/models?apps=llama.cpp&other=decision-model&sort=trending we also polished Llama App website & docs https://llama.app 🌟
JetBrains released Mellum2.1, a 12B mixture-of-experts model with 2.5B active parameters under the Apache 2.0 license, built for coding agents. Post-training shifted to reinforcement learning across thousands of environments and millions of sandboxed runs, and the model is available on Hugging Face. The source reports gains over Mellum2 on LiveCodeBench, AIME, GPQA Diamond, BFCL v4, IFEval, and SWE-bench Verified, and says it serves almost twice the tokens of Qwen3.5-9B under heavy load.
Why it matters: The post shows how reinforcement learning in real sandboxed environments changed a compact open model's repository work, with benchmark gains against Mellum2 and two peers.
Tencent Cloud has open-sourced Octop, a self-hosted multi-agent AI assistant platform aimed at families and small teams, with multi-user accounts and data kept on the user's own machine. The full text describes it as a single Python process that bundles the backend, web dashboard, CLI, IM gateway, cron jobs, and multi-agent runtime, with state rebuilt from SQLite on restart.
vLLM v0.31.0 is released with 717 commits from 307 contributors, including 96 first-time contributors. Highlights include DeepSeek-V4.1-Flash support, a vllm preload command that keeps weights in GPU memory across restarts, and Model Runner V2 with draft-model speculative decoding. The release also adds large-scale serving, scheduling, and HiSparse fixes, with full notes linked on GitHub.
Last week, several of South Korea's largest banks were hit by a cyberattack. A CrowdStrike report reportedly indicates the entire attack may have been carried out by one person. The attacker reportedly combined the open-source AI penetration tool ARTEX, DeepSeek v4.1-Flash, GLM-5.3, Grok 4.6, and Claude Code.
Huawei engineers presented XMFS, an experimental Linux kernel prototype filesystem, at the Linux Plumbers Conference in Prague on October 5. It aims to let applications reach cross-node shared memory on CXL 3.0 or Huawei unified bus servers through standard POSIX file calls. The code exists only on openEuler, not in the mainline Linux kernel.