Skip to contentSkip to stories
Updated

#Trend

Oct 7

  1. Google ResearchAI score62

    Google Research finds AI boosts patent drafting but junior lawyers' gains vanish without it

    AIA Google Research field experiment with 133 patent lawyers found AI tool access raised drafting scores by 0.34 to 0.38 standard deviations over three months. When the tool was removed for a redlining task, only senior lawyers kept an advantage of 0.45 SD, while junior lawyers showed no discernible improvement. The authors argue that tools which boost current output must not stop junior professionals from building the judgment that senior experts rely on.

    Why it matters: The field experiment separates AI's short-term productivity gains from skill retained after the tool is removed, which matters for training junior professionals.

Oct 4

  1. Epoch AIAI score62

    OpenAI researchers' coding-agent usage is doubling about monthly, Epoch AI reports

    AIOpenAI researchers' daily coding-agent usage, valued at API prices, rose from under $1 in January 2026 to $601 for the median researcher by mid-August. The 90th-percentile researcher reached over $7,000 per day, and both groups show doubling times of roughly one month. Epoch notes these are API-list values, not OpenAI's internal costs.

    Why it matters: The figures show internal coding-agent usage growing fast enough to matter for research cost, though they measure API-list value rather than OpenAI's actual spending.

Oct 2

  1. Epoch AI · The Epoch BriefAI score62

    Epoch AI estimates 2026 compute could run hundreds of millions of AI agents

    AIEpoch AI estimates that compute built from projected 2025 to 2027 high-bandwidth memory shipments could support tens to hundreds of millions of frontier AI agents, or billions of cheaper ones. Running nonstop, the top-tier agents would match the working hours of 140 million to 700 million full-time employees, and the central DeepSeek V4 Pro estimate of about 1.9 billion agents would match 8 billion workers.

    Why it matters: The estimate converts memory shipments into agent capacity and revenue ranges, showing how hardware supply could translate into labor and sales if demand keeps up.

Sep 30

  1. Anthropic ResearchAI score62

    Anthropic study finds robots can do most physical tasks but rarely cost-effectively

    AIAnthropic's research rates how well present-day robots can perform US job tasks, finding they can do 74% of physical tasks, or 34% of working hours, mostly in limited settings. Robots are cost-competitive for only 0.3% of job tasks, and at a 3% annual price decline it would take about 40 years to reach 10%. The report also finds robot-exposed jobs tend to pay less and be more physically demanding than LLM-exposed jobs.

    Why it matters: The report separates current robot capability from cost, showing that physical automation is technically broad but economically narrow for now.

Sep 29

  1. Exponential ViewAI score76

    Anthropic's S-1 shows revenue growing far faster than costs ahead of IPO

    AIAnthropic's draft S-1 prospectus, reported by Reuters, shows an $8bn operating loss and a $42bn net loss for 2025, which includes an accounting charge. The author argues revenues are growing far faster than costs, with the company likely turning a profit in 2026. The source also cites $518bn in compute commitments over 7-10 years, about 80% of which cannot be cancelled.

    Why it matters: The piece sets Anthropic's 2025 losses against its revenue growth and compute commitments, offering a concrete read on how an AI lab's finances could look at IPO.

Sep 25

  1. Kevin WeilAI score75

    Claude solves nine-loop scattering amplitude calculation past prior eight-loop record

    AIAnthropic reports that Claude solved a nine-loop calculation in the planar N=4 super-Yang-Mills model, surpassing the previous eight-loop record set by Lance Dixon and collaborators. The quoted post says Claude ran largely unsupervised for days in Claude Science using a single prompt, at a total cost of a few thousand dollars, and Dixon independently verified the result. Kevin Weil's own text praises the achievement and expects AI to advance high energy physics over the coming 12 months.

    Why it matters: The quoted Anthropic post gives a concrete benchmark: Claude ran for days to reach nine loops, extending the previous eight-loop record in a physics model.

Sep 24

  1. Google · Innovation & AIAI score62

    Google's Project Suncatcher will test TPUs in orbit on a prototype satellite

    AIGoogle's Project Suncatcher will launch a prototype satellite on the Transporter-18 rideshare mission with SpaceX to test how its TPUs handle spaceflight. Initial ground tests showed the Trillium TPUs survived vibration and a radiation dose greater than a five-year space mission would deliver. Google says cooling with heat pipes and radiators and laser links between satellites in 2027 remain open engineering challenges.

    Why it matters: The source reports concrete radiation, vibration, and cooling test results for TPUs, showing what space-based AI compute still has to solve.

Sep 23

  1. Microsoft ResearchAI score60

    Microsoft Research shows offloading robot AI inference improves performance and battery life

    AIMicrosoft Research reports that running physical AI inference on onboard GPUs can limit robot performance and battery life, while offloading inference to edge or cloud GPUs improved results in mobile manipulation tests. In its evaluation, smaller onboard GPUs slowed mapping and planning by up to 383% compared with an A100, and large onboard GPUs such as Jetson Thor drained robot batteries by up to 160%.

    Why it matters: The study measures how offloading robot inference to edge or cloud GPUs changes task success, battery life, and model size, offering evidence for infrastructure design.

Sep 22

  1. Noam BrownAI score78

    OpenAI releases GPT-6 Sol and Luna at 50% lower API prices

    AIOpenAI has released GPT-6 Sol and GPT-6 Luna, which it says build on GPT-6 Astra and offer faster, more affordable performance. API prices for Sol and Luna are 50% lower than GPT-5.6 promotional pricing, and Luna now costs $0.10 input and $0.50 output per 1M tokens. The author also notes an earlier 80% Luna price cut at the end of July, with output dropping from $6 to $0.50 within two months.

    Why it matters: The source gives concrete API price cuts across two model tiers, making the cost trend across recent releases easy to track for developers.

Sep 8

  1. Noam BrownAI score67

    OpenAI shares an AI-generated solution to the Navier-Stokes Millennium Prize Problem

    AIOpenAI says a group of agents using an unreleased next-generation model produced a solution to the Navier-Stokes Millennium Prize Problem, a question about whether smooth 3D fluid motion can break down that has stayed open for about 90 years. Noam Brown says the result cost millions of dollars, but argues that Astra now scores higher on ARC-AGI for about $20, versus roughly $500,000 for o3 on ARC-AGI 1.

    Why it matters: The post quotes OpenAI's claim about an AI-produced Navier-Stokes solution and adds cost comparisons that show how quickly test-time compute costs are falling.

Aug 27

  1. Epoch AI · The Epoch BriefAI score62

    Anthropic and OpenAI's 2026 revenue growth raises the question of how long it lasts

    AICombined annualized revenue for OpenAI and Anthropic reached $105 billion by August 2026, up 3.5 times from $30 billion at the start of the year. The author argues the key question is whether this growth comes from continued capability progress or from diffusion that will saturate. At the 3 times annual pace, frontier AI revenue would take about six years to reach today's world economy size.

    Why it matters: The piece tests whether OpenAI and Anthropic's hypergrowth reflects a temporary coding-agent spike or durable progress, using revenue scale to frame the question.

May 26

  1. MiniMax BlogAI score67

    MiniMax Agent Team Adds Parallel Multi-Agent Collaboration for Long Tasks

    AIMiniMax has upgraded its Agent, renamed Mavis, and introduced Agent Teams that run multiple role-based Agents in parallel on desktop. The team uses Leader, Worker, and Verifier roles so complex tasks can be split, checked, and reported at key checkpoints, and it merges TokenPlan and Agent Plan into one subscription with credits shared between Agent and API. The post also discusses the added token, handoff, and retry costs of multi-Agent work, and says the Agent will be open-sourced alongside MiniMax M3.

    Why it matters: The post explains why multi-Agent helps long tasks and where its verification, token, and aggregation costs come from, useful for judging when a team setup beats a single Agent.

Dec 4, 2025

  1. ARC PrizeAI score62

    ARC Prize 2025 results point to refinement loops as the central AI reasoning trend

    AIARC Prize reports that the top Kaggle entry reached 24% on the ARC-AGI-2 private dataset at $0.20 per task, and that all winning solutions and papers are open source. The top verified commercial model, Opus 4.5 (Thinking, 64k), scored 37.6% at $2.20 per task, while a Poetiq refinement on Gemini 3 Pro reached 54% at $30 per task. The author argues that refinement loops are the main driver of 2025 progress, and says ARC-AGI-3 is planned for early 2026.

    Why it matters: The post links 2025 competition results to a broader argument about refinement loops, showing how benchmark outcomes are being read as evidence of AI reasoning progress.

That’s everything