Skip to contentSkip to stories

Updated

Open source

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. Anthropic NewsroomOfficialAI score62

    Anthropic launches Cyber Mission with infrastructure defense and free OSS Scanner

    AIAnthropic has launched the Anthropic Cyber Mission, which starts with the Critical Infrastructure Defense Program for operational technology and OSS Scanner for open-source projects. The defense program brings frontier Claude models, on-site engineers and threat research to trusted providers such as Accenture, CrowdStrike and Palo Alto Networks. OSS Scanner gives enrolled open-source projects periodic free scans from its strongest models, with reports sent without human review and an expected true-positive rate above 90%.

    Why it matters: The announcement shows how a frontier AI lab is packaging cyber defense around critical infrastructure and open-source maintainers, including the program's partners and access routes.

  2. Anthropic ResearchOfficialAI score72

    Anthropic launches OSS Scanner, a free AI vulnerability scanner for open-source projects

    AIAnthropic is launching OSS Scanner, an opt-in service that runs periodic security scans of enrolled open-source projects using its strongest models at no cost. Its outputs are fully model-generated without human review, so some reports may be incorrect or invalid, though a pilot found 85 of 97 checked critical and high-severity findings met Anthropic's disclosure bar. Core maintainers of eligible projects can enroll through a GitHub pull request.

Oct 7

Oct 7Wed
  1. KhazixXAI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    AIOpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    Why it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. vLLMOfficialAI score46

    vLLM-Omni technical report unifies serving for omni-modality generation

    AIThe vLLM team released a technical report on vLLM-Omni, a unified serving runtime for omni-modality generation spanning multi-stage autoregressive pipelines, iterative diffusion, and stateful sessions. Current LLM servers and diffusion stacks each cover only one of these patterns, pushing deployments to stitch disjoint runtimes together. vLLM-Omni offers a shared control plane in which an orchestrator advances requests across stages, specialized engines handle compute, and a connector carries payloads.

    Image from @vllm_project's post
  3. meng shaoXAI score75

    Microsoft positions Windows as the home for hybrid AI agents across four layers

    AIMicrosoft has repositioned Windows as the home for hybrid intelligence, where AI agents can run locally or in the cloud. The announcement covers four layers: MXC reaching general availability for agent isolation, local models such as MAI Code 1.1 Flash, Copilot on Copilot+ PCs gaining local context and actions in coming months, and new hardware including RTX Spark PCs and DGX Station for Windows.

    Image from @shao__meng's post
  4. Orange AIXAI score34

    Next Token episode 5 covers Personal Agents, open-source software, and hardware projects

    AIThis Next Token episode discusses Personal Agents, including Dots in Codex, memory and cloud computer permissions, and whether agents should act as assistants or digital twins. The hosts also cover Instinct's booking and business-travel model, hands-on projects built with Opus 5.5, and whether software, games, and hardware could become open source as AI makes rewriting easier.

  5. MarkTechPostNewsAI score58

    Unsloth Studio re-checks changed model repos and blocks flagged weights before loading

    AIUnsloth Studio binds remote-code approval to a fingerprint of the scanned code, so changed code requires fresh consent before it runs. It also blocks weight files that Hugging Face has flagged for malware in the path the selected loader would deserialize. The article describes these checks as one layer among several, alongside package-content scans and OS sandboxes, and notes that the scanner is not a sandbox and cannot catch every evasion.

  6. InferactOfficialAI score38

    Inferact and partners cut vLLM TTFT nearly 70% at ~100K throughput

    AIInferact, working with DeepSeek, NVIDIA, and SemiAnalysis alongside the vLLM community, says joint work across models, custom kernels, and engine serving cuts time to first token (TTFT) by nearly 70% at ~100K throughput. vLLM is the open-source inference engine, and Inferact optimizes it for enterprise production deployments.

  7. Gizmodo · AINewsAI score51

    Vibe-Coded Artcraft Suite Offers Free Photoshop Alternative on GitHub

    AIDeveloper Brandon Thomas used Claude Opus 5.5 and Rust to build Artcraft, a free open-source suite with Photocraft, Vectorcraft, and other apps that mimic Adobe products. The author tested Photocraft and found its basic editing commands worked where expected, but Free Transform behaved unpredictably. Thomas describes the software as early alpha and invites developers to contribute.

  8. Google Developers BlogOfficialAI score62

    Google open-sources ML Drift, a cross-platform GPU engine for on-device AI

    AIGoogle's AI Edge Team open-sourced ML Drift under Apache 2.0, a GPU compute engine for on-device AI inference across OpenGL ES, OpenCL, Metal, and WebGPU. It serves as the core GPU acceleration engine within LiteRT and succeeds the legacy TFLite GPU delegate, which will no longer receive new features. The post cites benchmarks showing up to 40% lower frame latency in YouTube Shorts and up to 30% faster on-device performance in Adobe Lightroom and Photoshop.

    Why it matters: The post explains how ML Drift unifies GPU shaders across platforms and replaces the TFLite GPU delegate, which matters for developers deploying on-device models.

  9. Epoch AIOfficialAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  10. Hugging Face BlogOfficialAI score66

    How one developer built six custom models with ML-Intern for about USD 103

    AIA Hugging Face blog author used the ML-Intern agent in HuggingChat to build six small models by writing detailed prompts that specify datasets, base models, baselines, smoke tests, and spending limits. The projects include a citrus disease vision-language model, a Huggy character LoRA, a camera-angle LoRA, a doodle-to-object LoRA, a 0.8B prompt rewriter, and a 4-step distilled Agate model, with total compute cost of about USD 103. Each project's prompts and public models are linked from the post.

    Why it matters: The author shows how prompt structure, baselines, smoke tests, and budget caps shape an agent-driven training workflow, with per-project costs given.

  11. Teknium 🪽XAI score28

    Teknium posts a "hello" greeting on X

    AITeknium, a Nous Research affiliate, posted only the word "hello" on X. The quoted post from Nous Research announces a Series B raise to advance Hermes Agent and build a mobile app, with investors including NVIDIA and Samsung Next.

  12. Ars Technica · AINewsAI score46

    Artcraft releases open source clones of Adobe Photoshop, Premiere and other apps built with Claude

    AIDeveloper Brandon Thomas's Artcraft has launched seven open source apps in Rust that aim to replicate the interfaces and tools of Adobe Photoshop, Illustrator, Premiere, Lightroom, After Effects, InDesign, and Acrobat Pro. Thomas said he used Anthropic's Claude Opus 5.5 to build the clean-room replacements, with WebAssembly versions available for browser use. The apps remain in a "super early alpha" state, and commenters have pointed out many current shortcomings.

  13. Leandro von WerraXAI score36

    Snorkel expands Open Benchmarks Grants to $30M for AI evaluation

    AISnorkel AI is expanding its Open Benchmarks Grants tenfold to a $30M commitment to fund more diverse, robust, and continuously updated open AI benchmarks. The program adds an Open Benchmarks Red Team to test and strengthen those benchmarks, plus a Snorkel Research Fellowship for independent researchers developing new evaluation methods. The source says OBG-funded benchmarks have appeared on model cards from every major frontier lab.

  14. OpenRouterOfficialAI score46

    Perplexity Decider v1.1 arrives on OpenRouter with free output

    AIPerplexity's open-weights multimodal decision model, Decider v1.1, is now available on OpenRouter. It accepts text, JSON, or images and returns typed answers with probabilities, priced at $0.02 per million input tokens with free output. Perplexity says it scores highest on Hugging Face's new Decision Index 0.3 benchmark and costs half as much as v1.

  15. KushXAI score23

    Fluffles open-sourced as a lesson on stateful server-based agents

    AIDeveloper team open-sources fluffles-os on GitHub, presenting it as a hard lesson rather than the product at puffle.ai. The post says stateful server-hosted agents like Hermes proved unworkable, and the team's earlier fluffles agent, built as a near-unrestricted "god agent" on a Mac mini, was painful to harness because failure modes were unbounded. The team says it later ported to Eve, which let them focus on agent behavior instead of integration scaffolding, and plans a launch this week.

    Image from @kushbhuwalka's post
  16. GitHubOfficialAI score57

    GitHub Copilot local sandboxing becomes generally available

    AILocal sandboxing for GitHub Copilot is now generally available. It lets Copilot run commands in an isolated environment with controlled access to files, networks, system capabilities, and credentials. Enterprise teams can also centrally manage policies, and the feature is available in GitHub Copilot CLI, the GitHub Copilot app, and @code.

  17. TinkerOfficialAI score31

    IdeaLens detects whether ideas originated from humans or AI

    AIIdeaLens is a detector that identifies whether the ideas in a text came from a human or an AI, rather than judging the prose alone. On mixed-provenance benchmarks, it reached 81.3% average idea-detection accuracy, versus 25.4% for ProseLens and 25.9% for Pangram 4. The model, code, and data are open-sourced, and it was trained on Tinker.

  18. MarkTechPostNewsAI score60

    Liquid AI releases open-weight d1-3B and d1-omni-600M decision models

    AILiquid AI released two open-weight multimodal decision models, d1-3B and d1-omni-600M, which return probability answers in one forward pass with zero output tokens. d1-3B scores 48.57 on Decision Index v0.2.1 and answers one question in 8 ms on an RTX 4090, while the models are licensed free for commercial use below $10 million in annual revenue.

  19. OpenRouterOfficialAI score38

    Cloudflare's Clef decision models now available on OpenRouter

    AICloudflare's open-source Clef (27B) and Clef Flash (9B) decision models are available on OpenRouter. They accept text, JSON, or images and return typed answers with probabilities rather than generated text. Pricing is $0.24 per M input tokens for Clef and $0.09 per M for Clef Flash, with output free.

  20. Claude Code · GitHub ReleasesOfficialAI score36

    Claude Code v2.1.293 adds Claude Haiku 5.5 and fixes dozens of bugs

    AIClaude Code v2.1.293 adds Claude Haiku 5.5 (claude-haiku-5-5), now the default Haiku model on the Anthropic API, with 1M context and pricing of $0.10/$0.50 per Mtok ($0.50/$2.50 for prompts over 100K). The release also adds agentType to the subagentStatusLine payload and isDeferred to $.tool.register, and fixes numerous issues including a memory leak in HTTP MCP connections.

  21. a16z NewsBlogAI score46

    a16z backs Preference Model, which builds RL environments for training AI models

    AIPreference Model is open-sourcing Karotte, the framework it uses to build reinforcement learning environments that resist reward hacking, including defenses like killing stray processes before grading and rejecting grader-crashing files. The framework has been hardened through more than a million evaluation runs and controlled red-teaming. The company focuses on machine learning engineering tasks for leading labs, and a16z says it is partnering with Preference Model and its founders, Jennifer Zhou and Ning Cao.

  22. Georgi GerganovXAI score44

    llama.cpp adds ggml RPC for distributing inference across heterogeneous devices

    AIllama.cpp can distribute inference across heterogeneous devices through the ggml RPC backend, according to Georgi Gerganov. He says it is currently an advanced setting, but he expects it to become more accessible to regular users over time. A related post reports MiMo 2.6 Flash running across an RTX 6000 GPU and an M5 laptop over 10 GbE at about 40 tokens/sec.