Skip to contentSkip to stories

Updated

#Expert opinion

Oct 9

TodayOct 9Fri1 item
  1. QbitAIAI score62

    Google's AMIE Chatbot Tested in Real Pre-Visit Clinical Study Published in The Lancet

    AIA study led by Google and BIDMC tested Google's diagnostic AI chatbot AMIE with 98 outpatients before emergency visits, with a supervising doctor monitoring every exchange. No conversation needed interruption under the predefined safety criteria, and clinicians said AI summaries helped them prepare for 75% of visits. AMIE's differential diagnoses matched final diagnoses 90% of the time, but the authors say larger trials are needed.

Oct 8

Oct 8Thu
  1. Sherwin WuAI score62

    Harvey LAB-AA v1.1 adds hallucination gate, reshaping legal benchmark rankings

    AIArtificial Analysis and Harvey released LAB-AA v1.1, which credits a legal task only when deliverables pass every rubric criterion with no material hallucinations. Grok 4.7 (xhigh) leads at 9.4%, ahead of Muse Spark 1.3 (max) at 8.9% and GPT-6 Astra (max) at 8.6%, while over 60% of otherwise passing results contained a material hallucination. The sharper reordering appears in the hallucination counts, where GPT-6 Astra averages 0.03 material hallucinations per task against 13.96 for Gemini 3.8 Flash (high).

Oct 2

Oct 2Fri
  1. Redwood Research BlogAI score34

    Capabilities research pushes the safety-usefulness frontier too, not just safety research

    AIThe post argues that counting all research as safety work because it widens the safety-usefulness Pareto frontier is misleading. Safety research typically creates new safety options without boosting usefulness, while capabilities research typically raises usefulness at safety's expense, so developers tend to choose less safe points.

Oct 1

Oct 1Thu

Sep 28

Sep 28Mon
  1. Epoch AI · The Epoch BriefAI score62

    Epoch AI finds AI cost per benchmark score falling 13× per year

    AIEpoch AI estimates that the cheapest cost of reaching a given benchmark score has fallen about 13× per year over the past five years, faster than DNA sequencing, compute, lithium batteries, or electricity. Its example: a 75% GPQA Diamond score that cost about 30 cents per question with o3 in January 2025 cost $0.0004 per question with GPT-5.6 Luna under 18 months later. The authors caution that benchmarks are imperfect proxies for market prices, and the decline rate slows over time.

    Why it matters: The source compares AI price declines with other transformative technologies using benchmark-based cost estimates, giving readers a measured sense of how fast cost per capability is falling.

Sep 27

Sep 27Sun

Sep 25

Sep 25Fri
  1. AnthropicAI score78

    Claude solves a nine-loop scattering amplitude problem beyond the eight-loop record

    AIAnthropic reports that Claude solved a nine-loop scattering amplitude problem in planar N=4 super-Yang-Mills, surpassing the previous eight-loop record set by SLAC's Lance Dixon and collaborators. Working largely unsupervised for days from a single prompt, at a total cost of a few thousand dollars, Claude used methods developed by Dixon's group, and Dixon independently verified the result.

    Why it matters: The post shows Claude solving a nine-loop physics calculation beyond the previous eight-loop record, verified independently, which bears on AI use in theoretical physics research.

Sep 22

Sep 22Tue
  1. METR BlogAI score62

    METR's preliminary evaluation finds Claude Opus 5.5 is an incremental AI R&D gain over Fable 5.1

    AIMETR's preliminary evaluation concludes that Claude Opus 5.5 likely gives slightly higher AI R&D productivity uplift than Fable 5.1 but is unlikely to fully automate AI R&D. The evaluation used five capability tasks over 10 business days of API access, and METR says Anthropic reviewed and edited the summary before sign-off.

    Why it matters: The report separates two claims about AI R&D acceleration and discloses that Anthropic reviewed the summary, which helps readers weigh its independence and evidence.

Sep 9

Sep 9Wed
  1. Ahead of AI (Sebastian Raschka)AI score46

    GPT-6 Astra Leads Coding and Math Benchmarks, Shows Strong Computer Use

    AIOpenAI's GPT-6 Astra scores 99.9% on ARC-AGI-3, versus 7.8% for GPT-5.6 Sol, and leads Raschka's coding and math tests. Its strongest showing is in graphics and computer-use tasks, such as redrawing an image in a browser-based Paint app. The author notes that Artificial Analysis shows Astra at the frontier but not pulling far ahead on its Coding Agent Index.

Sep 4

Sep 4Fri
  1. Lewis TunstallAI score60

    Lewis Tunstall Shares Large Open Experiment on Autonomous Agents Iterating on NanoGPT Research

    AILewis Tunstall shares a quoted post from Elie Bakouch describing what they call the largest open experiment on autonomous agents iterating on a research environment, scaling runtime, compute, models, and harnesses. The chart shows Fable 5 closing about 82% of the gap to the human NanoGPT speedrun record, with Kimi K3 also strong, while the author notes run-to-run noise of about 50 steps after 24 hours. Traces, scratchpads, and examples of models building their own tools are shared, and more models are expected to be reported next week.

Aug 5

Aug 5Wed

Jul 3

Jul 3Fri
  1. Lil'Log (Lilian Weng)AI score62

    Lilian Weng surveys harness engineering as a path to recursive self-improvement

    AIThe post argues that the system surrounding a base model, called the harness, increasingly determines how well AI agents deploy and improve. It reviews research where harness components such as workflows, context, and code are optimized automatically through evolutionary search and meta-agent loops. The author concludes that evaluators, memory management, and human oversight remain open bottlenecks.