Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 1

Oct 1Thu
  1. Cloudflare Blog · AIOfficialAI score52

    Cloudflare AI Search reaches general availability with image embeddings and OCR

    AICloudflare's AI Search is now generally available, adding native image embeddings for visual retrieval, OCR for scanned PDFs, and a 10 MiB file limit up from 4 MiB. Billing begins November 1, 2026, charged per ingested token, stored GB-month, and query, with a free monthly allotment on all Workers plans.

  2. Cloudflare Blog · AIOfficialAI score62

    Cloudflare OS opens managed agent workspace waitlist with GitHub and Google Workspace support

    AICloudflare is opening a waitlist for fully managed Cloudflare OS deployments, where organizations configure a custom domain, Cloudflare Access policies, and an AI Gateway. The update lets agents mount existing GitHub repositories to explore code, fix bugs, and open pull requests, and read, draft, and send Gmail while accessing Google Drive. Built-in document, presentation, and spreadsheet tools can now export to Excel, CSV, PDF, Markdown, and HTML, with Word and PowerPoint export coming soon.

    Why it matters: The post shows how a managed agent workspace connects to GitHub and Google Workspace, which matters for teams weighing self-hosting against a managed deployment.

  3. Amazon ScienceOfficialAI score34

    Amazon Science Explains Graph-Centric Agentic AI for Network Root Cause Analysis

    AIAmazon Science describes a graph-centric approach in which a network digital twin graph and cascaded graph algorithms, orchestrated by an agentic AI layer, identify root causes in complex network failures. The approach was demonstrated with NTT DOCOMO at the Mobile World Conference, achieving root cause analysis in minutes on commercial networks. The article traces how graphs evolved from topology models to active reasoning substrates for agents.

  4. JetBrains AI BlogOfficialAI score75

    JetBrains Air enters early access as an agent system inside its IDEs

    AIJetBrains has opened the Early Access Program for Air, an agentic development experience available as a plugin on JetBrains Marketplace or in the 2026.3 EAP builds of its IDEs. Air works with existing agents such as Codex, GitHub Copilot, Junie, and Cursor, and it ships with no agents installed. Free Junie Lite runs are offered, while cloud runs require a JetBrains AI subscription.

    Why it matters: The post explains how Air brings existing agents into the IDE, showing a concrete workflow for managing parallel agent sessions alongside code review tools.

  5. WanOfficialAI score62

    Alibaba's Wan 3.0 ranks first overall on Artificial Analysis video leaderboard

    AIAlibaba's Wan 3.0 ranks #1 overall on the new Artificial Analysis AA-Video-T2V v2.0 text-to-video leaderboard, priced at $12 per minute of video. The benchmark judges models at 1080p using over 68,000 human preference votes across 1,000 prompts, and the author states Wan 3.0 leads 10 of 20 category boards.

  6. AI SupremacyBlogAI score50

    Google announces Gemini 4 Argon, its first frontier model since February

    AIGoogle announced Gemini 4 Argon, a model it says is built to sustain deep reasoning across complex, long-horizon workflows, roughly seven months after its last flagship release in February. The article says cybersecurity testing will be completed after October 1, with no benchmark scores, pricing, or availability details provided.

  7. Ai2 (Allen Institute for AI)OfficialAI score62

    Ai2 releases Olmo-core 3, an open framework for training large MoE models

    AIAi2 released Olmo-core 3, an open training framework redesigned to scale mixture-of-experts models into the trillion-parameter range. In one benchmark, expert count rose from 8 to 128 with about 3.2B active parameters per token, total capacity grew from 4.6B to 47B, and throughput fell by less than 5%. The framework is fully open, so researchers can train their own MoEs and experiment with routing and parallelism.

    Why it matters: The release documents concrete MoE scaling results and reported failure modes, useful for teams weighing training-stack tradeoffs before adopting an open framework.

  8. Ahmad Al-DahleXAI score10

    Meta's Ahmad Al-Dahle posts brief update announcing new product changes

    AIAhmad Al-Dahle, who leads Meta's Llama work, shared a short post saying the team has released new updates. The post itself gives no concrete features, versions, or figures, and it is quoting a post from Airbnb CEO Brian Chesky that lists trip-idea sharing with friends, AI search, and laundry and baby gear rentals.

  9. Anthropic ResearchOfficialAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

  10. Manus BlogOfficialAI score45

    Manus 2.0 Adds Video Editor for Creating and Editing Publishable Videos

    AIManus 2.0 introduces Video Editor, which lets users refine videos Manus generates, including changes to music, captions, and cut timing, without regenerating the entire video. The article describes Manus creating explainers, launch films, and animations from a single prompt, drawing on web search, video models such as Seedance 2.5, and code for motion graphics.

  11. Anthropic NewsroomOfficialAI score38

    Barclays expands Claude across operations, targeting 50% developer adoption by end-2026

    AIBarclays is expanding its collaboration with Anthropic to roll Claude out across its global operations, with Claude Code expected to reach 50% of its developer population by the end of 2026. Its Colleague Knowledge Assistant, powered by Claude through retrieval-augmented generation, has been used by more than 16,000 colleagues and handled over one million searches. In Global Markets, Claude models classify and route roughly 120,000 client emails daily.

  12. LangChain BlogOfficialAI score58

    LangChain shows how to build a model router in its Open SWE coding agent

    AILangChain built a model router inside its open source coding agent Open SWE that picks one of three models for each thread. In an A/B test against always using GPT-6 Astra, the median cost per thread fell 64% with no measurable change in merged PR rate. The router runs on the thread's first message, using a base prompt, per-tier criteria, and a classifier model, and the post lists next steps including subagent routing and mid-thread re-routing.

  13. Suno BlogOfficialAI score42

    Suno Launches Speech Beta, Generating Voice and Music Together as One Track

    AISuno has launched Speech in beta, which it describes as the first audio model that generates voice and music together as one cohesive track. Users type an idea or existing text, then describe the voice and musical style they want, and the beta is now open to everyone after a month of testing with a small group.

  14. Luma AI NewsOfficialAI score34

    Luma Launches Variants to Auto-Adapt Approved Static Ads Across Formats and Languages

    AILuma has launched Variants, which builds placement-ready versions of one approved static ad across five formats (Story 9:16, Portrait 4:5, Square 1:1, Medium Rectangle 6:5, Widescreen 16:9) and selected languages. Logos, headlines, and CTAs stay intact while layout and copy adapt to each placement and market. The first release covers static ads, resizing, and translation, and is available now from the Discover tab in Luma.

  15. Manus BlogOfficialAI score47

    Manus 2.0 Adds Game Dev for Building Multiplayer Games Without Coding

    AIManus has launched Game Dev in Manus 2.0, a feature that lets users with no coding experience build games with a real-time tweak panel, asset management, and multiplayer servers. The tweak panel lets users adjust settings such as speed, gravity, damage, and spawn rate while playing. Manus also handles much of the multiplayer infrastructure, including server deployment and networking, so games can be shared and played with friends.

Sep 30

Sep 30Wed
  1. Allie K. MillerXAI score33

    Jony Ive's OpenAI hardware rumors evolve from voice device to donut

    AISpeculation about Jony Ive's secret OpenAI hardware has shifted from a voice-based desk device in May 2024 to a wearable pin, then a pen, and now a 3D donut with a swivel piece. The post points to the Dots logo as a donut, suggesting the design has moved toward that shape. The author notes that the rumors remain unconfirmed.

  2. hardmaruXAI score26

    David Ha argues AI's future lies in orchestrating multiple models, not bigger ones

    AIIn a Nikkei Asia op-ed, Sakana AI co-founder and CEO David Ha argues that as frontier scaling slows, the next phase of AI will center on mastering orchestration rather than building ever-larger models. He contends that value shifts from individual model weights to intelligence that coordinates multiple models for different situations, and that sovereign AI rests on supply-chain resilience rather than national isolation.

  3. Sakana AIOfficialAI score33

    Sakana AI's David Ha argues the future of AI lies in orchestrators

    AISakana AI co-founder and CEO David Ha published a Nikkei Asia op-ed titled "The future of AI belongs to the orchestrators." He argues that ever-larger models face limits, as open models close the gap within months and frontier inference costs can exceed the hourly wage of the people they assist. He also contends that sovereignty means supply-chain strength, not national isolation.

  4. Guillermo RauchXAI score22

    Vercel AI Gateway rejects dubious token promos, prioritizing trustworthy providers and data privacy

    AIVercel says its AI Gateway turns down "free token" promotions from companies making dubious Zero Data Retention claims, prioritizing the best providers over provider count. The company argues it has the largest trustworthy view of global AI token flows, citing real usage from 400k+ paying customers, thousands of enterprises, and zero markup. It says it invests as heavily in legal, compliance, privacy, and back-office operations as in engineering to serve trillions of tokens daily.

  5. Hamel HusainXAI score38

    Hamel Husain Reviews Claude's New Auto Eval Plugin for Evaluations

    AIHamel Husain has published a longer review of a new Claude Auto Eval plugin after many users asked about it. He invites readers to share their experiences using the plugin and how it went for them. The plugin is part of Claude's ability to help build evaluations and hillclimb on them, as described by @ClaudeDevs.

    Image from @HamelHusain's post
  6. Max ZeffXAI score34

    Greg Brockman drops second $25M donation to AI super PAC

    AIOpenAI co-founder Greg Brockman is no longer making the second $25 million donation he had promised to the super PAC Leading the Future, according to a New York Times scoop. The commentary notes that AI regulation has become a major national issue far faster than many expected since Brockman's August 2025 commitment.

  7. indigoXAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

    Image from @indigox's post
  8. Apple Machine Learning ResearchOfficialAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  9. Apple Machine Learning ResearchOfficialAI score36

    RLTL;DR: Self-Improvement Through Internalized Self-Generated Feedback

    AIApple researchers introduced RLTL;DR, a reinforcement learning method in which an agent writes its own one-line insight after each failed attempt and learns to map tasks to those insights. On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at 0% to 1% Pass@1, while RLTL;DR reached 14–31% with insights in context and 12–13% without them at evaluation. A compact variant, SFTL;DR, trained on just 4k task-insight tuples recovered nearly the full performance of RLTL;DR.

  10. Latent.SpaceXAI score67

    OpenAI details agent stack with Computer Use, Dots, and Decisions API

    AILatent.Space shares a podcast episode in which OpenAI's AriX and Nikunj Handa discuss OpenAI's new agent stack. The episode covers Computer Use, Dots cloud computers for agents, and the Decisions API, which the speakers say went from idea to product in weeks.

    Video from @latentspacepod's post
  11. Google ResearchOfficialAI score40

    Google's science AI tops CDC flu hospital admission forecasts this season

    AIThe CDC announced that Google's science AI model ranked highest among 39 eligible models for forecasting flu-related hospital admissions during the 2025-26 flu season. Google's forecasts were built with Empirical Research Assistance, an AI tool that generates computational solutions across scientific fields.

    Image from @GoogleResearch's post
  12. Nathan LambertXAI score26

    Nathan Lambert welcomes Google's Gemini 4 as frontier competition

    AINathan Lambert says he is pleased to see Google surprise people with Gemini 4. He argues that more labs at the frontier benefit consumers through competition and reduce the concentration of power. He says he is eager to see how the model performs in real-world scenarios.

  13. Meta NewsroomOfficialAI score10

    Meta Names Dhruv Vohra Managing Director for Southeast Asia Business

    AIMeta has named Dhruv Vohra Managing Director, Global Business Group, Southeast Asia, to lead commercial strategy across Indonesia, Malaysia, the Philippines, Singapore, Thailand and Vietnam. He reports to Benjamin Joe, Vice President for Asia Pacific, and will oversee teams serving advertisers and agencies. Vohra joined Meta in 2019 and previously led the SMB Group across Asia Pacific.

Only the first 50 pages are available. Search or browse topics for older items.