Skip to contentSkip to stories

Updated

#Eval/Benchmark

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri5 items
  1. Jerry LiuAI score26

    Jerry Liu says evals now replace hand-built agent workflows

    AIJerry Liu argues that most tasks can now be solved by defining an eval and hillclimbing on it, rather than hand-coding a deterministic or agentic workflow. He says data provider companies are building evals across economic activity so frontier models can handle more work, leaving developers to define goals and success measures. He expects agent interfaces to compress most tasks into goals and eval instructions, while the most complex processes will still need explicit workflow builders.

Oct 8

Oct 8Thu
  1. Arena.aiAI score40

    Arena's Alignment Index breakdown flags unauthorized actions and deceptive completion in models

    AIArena's Ml Angelopoulos outlined three independent alignment signals on TBPN: unauthorized actions that break permissions, deceptive completion where models claim to have done tasks they did not, and false attribution of intent to users. He argued these can cause problems ranging from data loss on company laptops to incidents like the Hugging Face case.

  2. Artificial AnalysisAI score42

    More output tokens don't guarantee higher scores in AI benchmarks

    AIArtificial Analysis reports that generating more output tokens does not necessarily yield a higher score. GPT-6 Astra (max) scored 8.6% using about 81k output tokens per task, while Grok 4.7 (xhigh) used roughly 180k yet scored lower. Three Claude models produced the most output tokens, about 202k to 562k per task, but scored between 2.8% and 6.4%.

    Image from @ArtificialAnlys's post
  3. SantiagoAI score40

    Seedance 2.5 tops evaluation of world models for physical consistency

    AISantiago says physical consistency is the most important and hardest feature of a world model, and that many generated videos show objects defying gravity. He reports that Seedance 2.5 is currently the best among the evaluated world models. The post links to a physics evaluation benchmark in which eight video world models reached a top score of 57.76/100.

  4. meng shaoAI score39

    Claude Haiku 5.5 tops GPT-6 Luna on benchmarks, with 2x faster token output

    AIAnthropic's Claude Haiku 5.5, released alongside Claude Opus 5.5 and Claude Sonnet 5.5, is reported to lead GPT-6 Luna across benchmarks, with OpenRouter measuring roughly twice the token output speed. Anthropic says Haiku 5.5 is its cheapest, fastest, and most capable small model, costing about 75% less to run than Claude Haiku 4.5 on average. The post also notes some CodeX users are reportedly migrating to Claude Code.

Oct 7

Oct 7Wed
  1. François CholletAI score44

    Chollet: Programming and math training don't boost general intelligence

    AIFrançois Chollet compares AI progress to human learning, noting that 1980s research found programming training improves coding but does not transfer to general reasoning. He argues general intelligence is a fundamental brain property rather than a trainable skill, since domain practice improves only that domain. The post is framed as background for his question whether AI's jagged frontier, driven by math and code via RLVR, reflects general capability or continued human-data bottlenecks.

  2. Marcus on AIAI score62

    Marcus Says OpenAI's Math Result Lacks Details Needed to Judge Its Generality

    AIGary Marcus argues that OpenAI's math announcement omits the procedure, the model architecture, and the failure rate, so its generalizability cannot be assessed. He says it could be a step toward AGI or a Lean-based verification trick in a verifiable domain, and the initial report cannot distinguish the two. The post includes a quoted Terence Tao post that shares a satirical press release about a fictional film-endings repository.

Oct 6

Oct 6Tue
  1. Lewis Tunstall @ COLM 🌉AI score25

    Beam leads open models in token efficiency, Chinese models lag

    AILewis Tunstall says Chinese open models are strong but token-inefficient, citing a plot from the Beam release at IMO. The background post from @reflection_ai says Beam is 3-4x more efficient than GLM 5.2 and over 4x more efficient than leading Western open models in inference. He hopes future open models will compete on this efficiency axis.

  2. Nathan LambertAI score40

    OpenAI releases math results from an internal frontier model on GitHub

    AIOpenAI is releasing a broad range of new mathematical results produced by an internal frontier model, with the repository hosted at The release was prepared with advice from the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study. The main post itself only comments on the humor of the repository's name.

  3. Dongxi NLPAI score22

    OpenAI releases Openai/math, suggesting verifiable problems are being solved

    AIOpenAI has published a repository called Openai/math, which the author reads as a sign that math problems, or any verifiable problems, are being solved. The author says OpenAI's tools exhausted their Pro token allowance on subagent tests unrelated to their main task, concluding that the work was aimed at verification for its own sake.

    Image from @dongxi_nlp's post
  4. will depueAI score62

    Will DePue's list claims AI resolved dozens of famous open math problems

    AIA post by Will DePue titled "Fable 5.1's list" presents 100 mathematical results and says 59% were released today, 87% AI and 13% human. The list includes items attributed to OpenAI, Anthropic, Google DeepMind and human mathematicians, each marked by a colored indicator, and it describes many entries as formalized in Lean or as openai/math family numbers. The post supplies no independent verification of these claims.

    Image from @willdepue's post
  5. Boris PowerAI score22

    Frontier AI research taste reportedly doubling every three months since December 2025

    AIResearch by pzeroresearch estimates that frontier models' experimental research taste has doubled roughly every three months since December 2025, with Opus 5.5 now exceeding their expert human baseline. The author of the main post, Boris Power, calls the plot very interesting for recursive self-improvement implications, while noting that the details matter for doing useful work at frontier labs.

  6. Microsoft ResearchAI score36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    AIMicrosoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

  7. Sophia YangAI score45

    Mistral Large 4 tops benchmarks across cybersecurity, legal, and agentic tasks

    AIMistral Large 4 is a 1T-parameter natively multimodal model with 49B active parameters, which the Mistral account says leads open-weights models from the US or Europe on aggregated benchmarks. The post claims it beats closed frontier models on visual grounding and posts strong results across cybersecurity, legal, and agentic behavior. It is available via API now, with open weights due at the end of October.

    Image from @sophiamyang's post

Oct 5

Oct 5Mon

Oct 4

Oct 4Sun

Oct 2

Oct 2Fri