Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 9

TodayOct 9Fri
  1. elvisXAI score60

    Meta researchers propose agent plasticity to measure self-improvement efficiency

    AIResearchers from UC Berkeley, Meta Superintelligence Labs, and other institutions introduce agent plasticity, the gain on held-out tasks per dollar of learning cost, with model weights frozen. The paper reports that in chess, Go, and Hex, Claude Fable 5 reaches the highest final score while GPT-5.6 Sol gains the most per dollar, and in NetHack only Claude Opus 5.5 improves significantly.

    Image from @omarsar0's post
  2. 404 MediaNewsAI score11

    Behind the Blog: 404 Media discusses AI and spirituality

    AI404 Media's Behind the Blog column discusses AI and spirituality, with Jason saying the outlet writes about AI's current capabilities and harms rather than dismissing it outright. He says reporters sometimes test AI tools while working on stories to write from an informed perspective. The excerpt does not say more about the spirituality discussion.

  3. TechCrunch · AINewsAI score44

    a16z's Olivia Moore says consumer AI revenue is mostly prosumer and many categories lack AI apps

    AIAndreessen Horowitz partner Olivia Moore released a report on the top 100 consumer AI apps, finding ChatGPT still leads by a wide margin while smaller players like Suno and ElevenLabs show staying power. Moore says almost all AI revenue comes from subscriptions and token usage, and that most consumer AI is prosumer AI. The report finds no top-100 entrants in social, dating, marketplace, retail, travel, finance, or health categories.

  4. Mike KnoopXAI score62

    Tufa Labs hits 88.06% on ARC-AGI-2, clearing the Kaggle bonus threshold

    AIMike Knoop says the 85% Grand Prize bonus threshold has been reached on Kaggle. The ARC Prize 2026 leaderboard lists Tufa Labs first at 88.06%, followed by Rabbithole at 80.56% and Yi-Chia Chen at 77.22%. Knoop says this will be the final year for ARC-AGI-2 on Kaggle and expects an open-source, low-cost, offline reproducible solution and model.

  5. Ethan MollickXAI score23

    Google's post-Gemini 4 challenge is product integration, Mollick argues

    AIEthan Mollick says Google's main challenge after Gemini 4 is what it does with a strong model. He argues that Anthropic and OpenAI are moving toward a single interface for many tasks using orchestrator agents. He says the fragmented products of the Gemini 3 era will not work for what comes next.

  6. Ai2OfficialAI score22

    Ai2 replaces its GPU scheduler after idle jobs hoarded capacity

    AIAi2 says its old scheduler made every scheduled workload eventually run at HIGH priority. Researchers kept idle jobs running to reserve GPUs for experiments, because the incentives rewarded holding capacity even with no active work.

  7. Ai2OfficialAI score13

    Ai2 describes a fair-share GPU scheduler for research budgets

    AIAi2 says managers now assign GPU-time budgets to research programs and projects. Its fair-share scheduler compares recent usage with those budgets and moves work from underused allocations ahead in the queue.

    Image from @allen_ai's post
  8. Ai2OfficialAI score25

    Ai2's new scheduler delivers 98% of owed GPU hours in 30-day test

    AIAi2 reports that over a 30-day test of its new scheduler, teams received 98% of the GPU hours they were owed, based on actual demand. Cluster occupancy stayed at 98%, and spare capacity went to interruptible work without drawing down team budgets.

  9. The Algorithmic BridgeBlogAI score40

    Meta's AI comeback follows heavy Anthropic Claude spending and a new Muse Spark model

    AIMeta spent heavily on Anthropic's Claude models, with internal use reaching up to 60,000 employees and a projected $10 billion yearly spend, according to The Algorithmic Bridge. The author says Meta then released Muse Spark, which scored 52 on the Artificial Analysis intelligence benchmark, on par with Claude Opus 4.6.

  10. Tessl BlogOfficialAI score36

    Tessl's agentic code review splits PR checks into standards, lenses, and memory

    AITessl Blog describes an agentic code review workflow built for teams whose coding agents produce pull requests faster than humans can review them. The workflow runs review against a written standard in the repository, applies four parallel perspectives covering correctness, security and privacy, scale and resilience, and maintainability, then records each finding, verdict, and response. Tessl Code Review, which the post says is free to start, runs these perspectives as skills, and the team's memory of past decisions is fed back into the standard.

  11. Ai2 (Allen Institute for AI)OfficialAI score46

    Ai2 describes GPU time budgets that replaced its priority-based cluster scheduler

    AIAi2's AI Infrastructure team replaced its priority-based scheduler for GPU clusters with GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change moved debates over how much GPU time each research project deserves from case-by-case operational decisions into a transparent budgeting process. The clusters range from 88 to 1024 GPUs across NVIDIA H100, B200, and B300 hardware, and serve about 150 internal researchers.

  12. AWS Machine Learning BlogOfficialAI score67

    How Postman runs Agent Mode for 40 million developers on Amazon Bedrock

    AIPostman describes the architecture behind Agent Mode, its AI agent for API testing, documentation, discovery, and implementation. The post covers limiting tools per task, using schema-based queries, building purpose-shaped context handlers, and running on Amazon Bedrock with cross-Region inference and prompt caching. Postman reports that tool-selection errors rose once the visible toolset exceeded about 40 tools.

    Why it matters: The post shows concrete patterns for tool scoping, context handling, and Bedrock routing and caching, which apply to any team moving an agent past a prototype.

  13. AWS Machine Learning BlogOfficialAI score36

    AWS recaps September 2026 Bedrock, AgentCore, and Strands updates for AI builders

    AIAmazon Bedrock Managed Agents, powered by OpenAI, entered public preview, and OpenAI's GPT-6 Astra, GPT-6.1 Sol, and GPT-6.1 Luna became generally available on Amazon Bedrock. AWS also released Strands Decider 2B, a 2B-parameter open source decision model that answers in about 115ms locally, and said the Strands harness uses 28 percent fewer tokens than popular harnesses while matching their accuracy.

  14. Hugging Face BlogOfficialAI score38

    Ai2 replaces priority scheduler with GPU time budgets for cluster allocation

    AIAi2's AI Infrastructure team replaced its priority-based GPU cluster scheduler with a system using GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change turns decisions about how much GPU time each research project receives into a transparent administrative budgeting process. Its clusters, which range from 88 to 1024 GPUs including H100, B200, and B300 units, serve about 150 researchers facing demand two to three times available capacity.

  15. ElevenLabsOfficialAI score32

    ElevenLabs partners with Banner Health on AI voice agents for patient calls

    AIElevenLabs says it is partnering with Banner Health to answer patient calls with AI voice agents, starting with primary care scheduling. The ElevenAgents system books, reschedules, or cancels appointments directly in Banner's electronic medical record at any hour, and transfers calls to a Banner team member with context when a patient asks for a person.

    Image from @ElevenLabs's post
  16. elvisXAI score40

    Syren Video learns your style to build AI videos from prompts

    AISyren Video, a new agentic video tool, learns preferred graphics, motion, and editing rhythm from a user's library and generates new videos from a prompt. Users refine the results through chat, and the tool is free to try in a browser or through Claude MCP, per the company's announcement. The post's author says the education sector is exploring it.

  17. ARC PrizeOfficialAI score42

    ARC Prize 2026 ARC-AGI-2 high score reaches 88.06%

    AITufa Labs posted an 88.06% score on ARC-AGI-2, a new high for the ARC Prize 2026 leaderboard. ARC Prize says a $150K bonus prize, on top of guaranteed prizes, will be split among all teams scoring over 85%.

    Image from @arcprize's post
  18. Ben TossellXAI score10

    Ben Tossell lists AI meeting tools that alert and take notes

    AIBen Tossell lists seven AI meeting tools, including dot, Grok bot, Poke, Instinct, OpenAI's meetings plugin, Granola, and Dia, that alert users to meetings, join Zoom calls, or take notes. He says the list means he should no longer miss meetings, but he says he was late because of the post itself.

  19. Andrew CurranXAI score62

    OpenAI responds to three fired employees' letter on safety and trust

    AIOpenAI's research leaders say they parted ways with Jasmine, Mikita, and Tomek after an investigation found they violated policies on handling sensitive information. The company says the decision was not about raising safety concerns and that it is finalizing contracts with third-party safety assessors, with details to follow in the coming weeks.

  20. Hamel HusainXAI score22

    How to make the team case for investing in evaluations

    AIHamel Husain advises reviewing real user interactions and showing the team the problems found. He says to then fix those problems and show what improved as the case for investing in evaluations.

    Image from @HamelHusain's post
  21. Rohan PaulXAI score46

    Sabi raises $50M to build AI brain-reading baseball cap

    AISabi has raised a $50M seed round led by Vinod Khosla and Accel to build a baseball cap that turns brain signals into AI prompts. The cap can hold up to 100,000 sensors of 1 to 5 millimeters each, feeding a brain foundation model trained on 100,000 hours of labeled neural recordings. Sabi says the cap can predict a user's next 3-4 keystrokes from brain signals before they type them.

    Image from @rohanpaul_ai's post
  22. 🚨 AI News | TestingCatalogXAI score62

    Sabi raises $50M for a wearable brain-reading cap

    AISabi has raised $50M from Khosla Ventures, Accel, Initialized, DST, and Collab Fund, with Kevin Weil also participating. The company is building the Sabi Cap, a non-invasive wearable that aims to read brain signals and turn thoughts into text. Sabi says it designs its own custom chips and neuroimaging sensors, including a non-contact EEG chip fabricated by TSMC.

    Image from @testingcatalog's post
  23. SantiagoXAI score43

    Sabi raises $50 million for a wearable brain-to-text cap

    AISabi has raised $50 million from Khosla Ventures, Accel, Initialized, Kevin Weil, and DST Global to build a wearable brain-computer interface. The company says its TSMC-fabricated chip reads brain signals without touching the scalp, with custom sensors collecting data and its own AI model decoding it into text. The device is a cap rather than an implant.

  24. PixVerseOfficialAI score14

    PixVerse hosts sessions demoing ChatGPT plugin video creation

    AIPixVerse says each session includes a platform walkthrough, a live OpenAI demo showing ChatGPT generating a creative brief and finished video via the PixVerse Plugin, a creator sharing their workflow, and live Q&A. The post presents these as recurring sessions rather than a new product launch.

  25. OpenAI DevelopersOfficialAI score46

    Codex on Windows gets new MXC-based sandbox mode

    AIOpenAI says Codex on Windows now has a new sandbox mode built on Microsoft's Execution Containers (MXC), offering faster setup, stronger network enforcement, and granular file access controls. The mode requires a compatible Windows 11 device. Background from Microsoft's announcement says MXC is now generally available on Windows 11, keeping agents within boundaries the operating system enforces.