Skip to contentSkip to stories

Updated

Open source

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. Claude Code · GitHub ReleasesOfficialAI score56

    Claude Code v2.1.295 adds hook failure blocking and gateway controls

    AIClaude Code v2.1.295 adds onFailure: "block" for command and HTTP hooks, so a hook that cannot start, times out, or exits unexpectedly blocks the action. The release also adds an optional models list for Claude apps gateway upstreams, plus upstream_request_id in the inference audit event, and fixes a range of MCP, plugin, and terminal issues.

  2. Sherwin WuXAI score60

    Harvey LAB-AA v1.1 adds hallucination gate; Grok 4.7 leads at 9.4%

    AISherwin Wu, an OpenAI employee, says the updated Harvey LAB-AA v1.1 benchmark, announced by Artificial Analysis with Harvey, is more useful than the original LAB results. The new Hallucination-Gated All-Pass Rate credits a task only when every rubric criterion passes and no material hallucination appears. Grok 4.7 (xhigh) leads at 9.4%, while GPT-6 Astra (max) at 8.6% has very few material hallucinations.

    Why it matters: The update adds a hallucination gate to a legal benchmark, showing that models with high all-pass rates can rank much lower once material errors count.

  3. Codex · GitHub ReleasesOfficialAI score36

    Codex 0.162.0 adds managed worktree tools and clickable URLs in the TUI

    AIOpenAI's Codex 0.162.0 release adds tools for creating and listing managed Git worktrees from trusted local projects when the worktrees feature is enabled. The update also lets users pin tasks in the agent Command Center, copy transcript blocks with /copy, and make URLs clickable in approval headers, questions, and warnings, along with several Linux and Windows sandbox fixes.

  4. Tessl BlogOfficialAI score42

    Agent Skills Should Be Treated as Supply Chain Components

    AITessl's talk at AI Native DevCon London argues that agent skills, which can be markdown files with instructions and bundled material, act as supply chain components that can shape agent behavior. The author says reading SKILL.md once is insufficient because risks can sit in supporting files, updates, and workspace trust settings. He identifies the danger as the combination of private context, untrusted content, and external communication, and cites research scanning roughly 4,000 public skills for issues including malware-like behavior.

  5. elvisXAI score42

    Voyager: an open harness for creative AI work across video and games

    AIElvis Saravia argues that creative work needs domain-specific agent harnesses rather than coding-oriented ones, and he highlights Voyager as an open harness for video, graphics, and games. According to the quoted post, Voyager lets agents work with local files and drive apps such as Blender, DaVinci Resolve, and Unity, and it is designed to work with models like Opus, Astra, and DeepSeek.

    Video from @omarsar0's post
  6. Artificial AnalysisOfficialAI score28

    Artificial Analysis Pareto frontier: GPT-6 Luna cheapest per task at $0.22

    AIAmong models with a Hallucination-Gated All-Pass Rate above 0%, GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max), and Grok 4.7 (xhigh) set the Pareto frontier for score versus cost per task. GPT-6 Luna (max) is the cheapest at about $0.22 per task, scoring 3.3%, while Grok 4.7 (xhigh) leads at about $9.50 per task and Muse Spark 1.3 (max) costs about $4.20. The three Claude models cost about $18 to $22 per task.

    Image from @ArtificialAnlys's post
  7. DatabricksOfficialAI score32

    Databricks' Vibe Data Modeling builds business-specific data models with an agent

    AIDatabricks introduced Vibe Data Modeling, an open-source agent that helps teams build, validate, and evolve business-specific data models. It applies roughly 250 modeling rules while keeping data modelers and business stakeholders involved. Teams can start from 40 industry models as a baseline and iterate toward models that reflect how their business operates.

    Video from @databricks's post
  8. laurenXAI score29

    Omarchy seeks feedback on Grok Bot plugins and integrations

    AILauren Tan invites users of Grok Bot on Omarchy and developers building plugins for it to share feedback and feature requests. The post points to the Omarchy plugin catalog and asks what integrations could be supported. Background from DHH says SpaceXAI joined the Omacom Foundation as a Founding Corporate Patron, contributing $1,500,000 in Grok tokens for Omarchy's maintenance and development.

  9. PyTorch BlogOfficialAI score62

    NVIDIA Dynamo adds session-level IDs to route and cache agentic inference

    AINVIDIA Dynamo uses a unified session-level identifier to make its inference stack aware of agent sessions, subagents, and their KV cache across turns and tool calls. On SWE-bench, two TP4 MiniMax-M2 replicas on one 8xH100 node gained roughly 12-16% throughput from program-aware scheduling over KV-aware routing alone. The post also describes experimental shared-pool indexing and a proposed KvHint interface for session-aware cache policies in vLLM and SGLang.

    Why it matters: The post explains how session identifiers let an inference stack track agent working sets, with measured throughput gains on SWE-bench and agentic RL rollouts.

  10. KalaXAI score34

    Mistral Large 4 and Reflection Beam promise open weights this month

    AIMistral Large 4 and Reflection Beam are previewed now, with Mistral saying weights drop at the end of October and Reflection promising Apache 2.0 weights this month. The post argues that these announced future weights should be treated as a conditional migration dependency, not a current self-hosting option. API previews can be trialed immediately, but they do not prove an unreleased checkpoint will behave the same when downloaded.

  11. The Verge · AINewsAI score30

    SpaceXAI Backs Omarchy Linux Distro With $1.5 Million in Grok Tokens

    AISpaceXAI is joining the Omacom Foundation, which oversees the Omarchy Linux distribution, as a Founding Corporate Patron and donating $1.5 million worth of Grok tokens to the project. According to David Heinemeier Hansson's blog post, the tokens will primarily accelerate development, review code, and patch bugs. The partnership follows earlier controversy over Hansson's anti-immigration posts, which have drawn criticism of Omarchy's corporate contributors, including 1Password and Cloudflare.

  12. Alexander DoriaXAI score46

    LightOnOCR-3 claims state-of-the-art OCR performance under 1B parameters

    AILightOn has released LightOnOCR-3, a family of OCR models in 0.8B and 4B versions that it says lead benchmarks including OlmOCR-Bench and ParseBench, with the 0.8B model positioned as the sub-1B option. The models recognize text, handwriting, images, charts and document structure in one pass, process documents up to twice as fast as LightOnOCR-2, and are released under the Apache 2.0 license.

    Image from @Dorialexander's post
  13. Dhravya ShahXAI score42

    MemoryRepo: open-source implementation of Cognition's dreaming agent memory

    AISupermemory introduces MemoryRepo.dev, an open-source implementation of Cognition's dreaming memory system built on Cloudflare Artifacts, Durable Objects, Alchemy, and Effect. The project follows Cognition's Devin memory design, which builds a memory graph across sessions and prunes stale records overnight. Supermemory says it will incorporate learnings from this research into its own product.

    Video from @DhravyaShah's post
  14. ClineOfficialAI score46

    Cline makes Step 5 Preview free, citing strong DeepSWE coding scores

    AICline says Step 5 Preview is now free in its coding tool and scores ahead of Kimi K3 and GLM-5.3 on DeepSWE. The company describes it as one of the strongest open-weights coding models available. StepFun's background announcement describes Step 5 Preview as a 600B total / 27B active MoE model with 1M context and vision, and says open weights arrive on Oct 15.

    Image from @cline's post
  15. Andrew CurranXAI score13

    Andrew Curran Posts "The saga continues" Amid Tightened κ Result

    AIAndrew Curran posted a brief "The saga continues" update, with no clear publisher or model identified. Quoted context from @0xdoug reports a validated, merged PR that tightened κ from 2⁻¹⁸² to 2⁻¹⁵, described as a 500-thousand-fold improvement over the previous result and a 2^167-fold improvement over the original OpenAI result. The quoted post credits a community effort and says results are being verified and published.

    Image from @AndrewCurran_'s post
  16. Tessl BlogOfficialAI score44

    Continuous AI Brings Agentic Automation to Repository Workflows

    AITessl's blog post argues that repository automation needs Continuous AI, a third pillar alongside CI and CD for scheduled, auditable AI workflows that improve repositories over time. The article describes GitHub Agentic Workflows, which harden agentic workflow specifications into GitHub Actions that can run coding agents such as Claude Code, Copilot CLI, Gemini CLI, or Codex-style agents. It emphasizes read-only agent steps, restricted outputs, and human review of pull requests.

  17. elvisXAI score46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    AIRSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  18. Daniel HanXAI score38

    Unsloth adds OS-level sandboxing for Linux, Mac, and Windows

    AIUnsloth now supports OS-level sandboxing on Linux via bwrap, on Mac via seatbelt, and on Windows via Microsoft's MXC. Per-tool-call latency is under 100ms across all three, and its software-style sandboxing with regex AST checks adds about 3ms. The Windows integration was built in collaboration with Microsoft.

  19. MarkTechPostNewsAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  20. Goodfire ResearchOfficialAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  21. Leandro von WerraXAI score70

    Carbon-A open model and database predict 566 million gene candidates across 22,617 species

    AICarbon-A is an open model that predicts gene locations directly from DNA, and it has been used to annotate genomes from over 22,000 species. The release includes a database of 566 million gene candidates, about 16 times the gene annotations in the RefSeq dataset. Wet-lab RNA experiments supported 239 candidates missing from RefSeq across cats, Syrian hamsters, chickens, and Arabidopsis.

    Why it matters: The source ties an open gene-annotation model to specific wet-lab checks and gene counts, helping readers judge how far its predictions extend beyond well-studied genomes.

  22. Thomas WolfXAI score67

    Carbon-A open model and database find 566 million candidate genes across 22,617 species

    AIThomas Wolf says Carbon-A, an open model that finds genes directly in DNA, has been released with a database of 566.34 million candidate genes across 22,617 species. The team reports wet-lab validation of several new genes in cats, chickens and arabidopsis, and RNA evidence for 239 genes missing from reference annotations of common species.

    This story has a top pick“Carbon-A open model and database predict 566 million gene candidates across 22,617 species”

  23. Philipp SchmidXAI score46

    SynthID Detector now publicly available for verifying AI-generated content

    AIGoogle's SynthID Detector is now publicly available, letting users check whether an image, video, or audio file was generated by supported tools. Per the post, it scans for watermarks from Google and partners, including Nano Banana 2.1, OpenAI, NVIDIA, and Kakao, with Apple support coming soon. Uploaded files are deleted right after scanning.

    Video from @_philschmid's post
  24. Karl's AI WattsXAI score14

    Claude Opus 5.5 gains traction for weekly product videos and GoodCase expansion

    AIThe author says Opus 5.5 keeps improving and works well for producing weekly product short videos, with all materials generated directly without extra services. GoodCase added 269 new AI showcase cases, prompts, and 7 new Skills, bringing its total to 1,699 cases, 95 Skills, and 426 creators. The post also highlights awesome-seedance, which now lists 795 video cases, 367 prompt retests, 27 prompt templates, and 77 installable video Skills.

    Video from @aiwarts's post