Skip to contentSkip to stories

Updated

#Eval/Benchmark

Showing low-relevance items too. Hide low-relevance items

Oct 6

Oct 6Tue
  1. Epoch AIOfficialAI score60

    Epoch AI finds frontier models fall short of an end-to-end AI research task

    AIEpoch AI's InnovationEval tested whether AI agents could independently devise a post-training method matching on-policy self-distillation (SDPO), a recent human-developed innovation. GPT-5.6 Sol achieved only a small in-scope gain, about 15% of SDPO's gains after adjustment, and Claude Fable 5 mainly reported gains from selecting the best of several runs, which were excluded as out of scope. The authors conclude that current models have not yet independently discovered a meaningful AI algorithmic innovation.

    Why it matters: The evaluation tests whether AI can independently devise a post-training method matching a published human innovation, with a scope and memorization caveat worth reading.

  2. Dongxi NLPXAI score22

    OpenAI releases Openai/math, suggesting verifiable problems are being solved

    AIOpenAI has published a repository called Openai/math, which the author reads as a sign that math problems, or any verifiable problems, are being solved. The author says OpenAI's tools exhausted their Pro token allowance on subagent tests unrelated to their main task, concluding that the work was aimed at verification for its own sake.

    Image from @dongxi_nlp's post
  3. Thomas WolfXAI score38

    OpenAI releases new mathematical results from internal frontier model

    AIOpenAI is releasing a broad range of new mathematical results produced by an internal frontier model, with release guidance from the Institute for Advanced Study's Advisory Group on Mathematics and Artificial Intelligence. The results are available on GitHub at openai/math. The post itself is brief and emphasizes the results rather than hype.

  4. will depueXAI score62

    Will DePue's list claims AI resolved dozens of famous open math problems

    AIA post by Will DePue titled "Fable 5.1's list" presents 100 mathematical results and says 59% were released today, 87% AI and 13% human. The list includes items attributed to OpenAI, Anthropic, Google DeepMind and human mathematicians, each marked by a colored indicator, and it describes many entries as formalized in Lean or as openai/math family numbers. The post supplies no independent verification of these claims.

    Image from @willdepue's post
  5. Teknium 🪽XAI score33

    Hermes Index launches to rank models for Hermes Agent users

    AITeknium announced Hermes Index, which combines scores from the new HermesBench and three other agent benchmarks. The index aims to help Hermes Agent users find the best model at a given time and at a given price point. It was introduced by Nous Research as a way to inform model choice and show labs their performance in Hermes.

  6. Nous ResearchOfficialAI score34

    Claude Opus 5.5 tops new benchmark at 63.31 per-task score

    AIClaude Opus 5.5 leads the benchmark with a score of 63.31 at $4.99 per task, ahead of GPT 6 Astra at 56.25 ($11.61) and Sonnet 5.5 at 53.14 ($2.82). At the low end, DeepSeek V4.1 Flash scores 36.91 at $0.26, and Ling 3.0 Flash scores 21.56 at $0.054.

    Image from @NousResearch's post
  7. MIT News · AIOfficialAI score23

    MIT Lincoln Lab's LAICS Survey Tracks AI Accelerator Performance and Power Trends

    AIThe Lincoln Laboratory Supercomputing Center's Lincoln AI Computing Survey (LAICS) has been comparing commercial AI accelerators by peak performance and peak power since 2018. The latest paper covers more than 120 accelerators, up from 57 in the first, with data drawn from public sources. The team says five to 10 new AI accelerator startups emerge each year, and six have announced their first accelerators in recent months.

  8. Aravind SrinivasXAI score40

    Perplexity halves Decision API input pricing to $0.02 per million tokens

    AIPerplexity cut its Decision API input pricing by half, to $0.02 per million input tokens. The reduction follows the release of pplx-decider-v1.1-27b, an open-weights multimodal decision model that scores highest on Hugging Face's Decision Index 0.3 benchmark. The model costs half as much as v1.

  9. 👩‍💻 Paige BaileyXAI score5

    Paige Bailey thanks Stanford AI Lab, DeepMind, and Surge for event turnout

    AIGoogle/Gemini's Paige Bailey thanked attendees for supporting an event featuring Stanford AI Lab, Google DeepMind, Surge AI, and Hands-On Robotics. A companion post from Ugur Yektabasak praised a panel with Bailey, Amaub, and Teresa Nguyen during Tech Week and credited Google DeepMind with organizing it.

  10. elvisXAI score22

    Elvis Saravia urges learning to build good evals for domain edge

    AIElvis Saravia argues that building good evals on top of AI systems can put a practitioner at the frontier of their domain or task quickly. He advises readers to learn eval construction, calling it worth the investment. The post responds to Garry Tan's point that agents writing markdown skills on cron jobs can handle most knowledge work.

  11. Boris PowerXAI score22

    Frontier AI research taste reportedly doubling every three months since December 2025

    AIResearch by pzeroresearch estimates that frontier models' experimental research taste has doubled roughly every three months since December 2025, with Opus 5.5 now exceeding their expert human baseline. The author of the main post, Boris Power, calls the plot very interesting for recursive self-improvement implications, while noting that the details matter for doing useful work at frontier labs.

  12. elvisXAI score41

    Parsewave audit fixes 206 verifier bugs in AutomationBench

    AIParsewave audited all 600 public tasks in Zapier's AutomationBench and human review confirmed 206 real verifier bugs, all of which were fixed in AutomationBench Verified. Replaying 1,235 Kimi K3 runs on the old and fixed verifiers changed 27.9% of grades, with pass rate rising from 18.8% to 43.8% where verifiers were too strict and falling from 60.2% to 49.7% where they were too lenient.

  13. Yuchen JinXAI score34

    Reflection's Beam and Mistral Large 4 near GLM-5.2 level

    AIYuchen Jin says Reflection's Beam and Mistral Large 4 both reached roughly GLM-5.2 level within the past two days. He suggests the Western versus Chinese open-source model gap may come down to Chinese labs being able to distill Anthropic and OpenAI models, which Western labs cannot.

  14. Microsoft ResearchOfficialAI score16

    Jennifer Neville on winding research paths and practical AI evaluations

    AIIn a Microsoft Research Podcast episode, Jennifer Neville discusses her nonlinear route into computer science and her push for more practical evaluations of today's AI systems. The post offers little beyond this framing, so no specific models, benchmarks, or results are mentioned.

    Video from @MSFTResearch's post
  15. ARC PrizeOfficialAI score39

    Grok 4.7 reasoning tokens track its ARC-AGI-2 public scores

    AIARC Prize reports that Grok 4.7's reasoning-token usage correlates with its ARC-AGI-2 public scores. The low setting averaged about 10k tokens per test-pair attempt and scored 25%, while medium through xhigh used 86k to 120k tokens and scored 57.5% to 60%. ARC Prize suggests the lower token usage may help explain the low setting's lower score.

    Image from @arcprize's post
  16. ARC PrizeOfficialAI score28

    Grok 4.7 scores 1.8% on ARC-AGI-3 standard harness

    AIGrok 4.7 scored 1.8% on ARC-AGI-3 in the standard harness, which lets models carry notes between turns, slightly below the 2.1% reported for Grok 4.7 in that setting. In a new provider adapter harness that preserves opaque reasoning and enables auto compaction, the score rose to 10.0%.

  17. ARC PrizeOfficialAI score22

    Grok 4.7 uses more reasoning tokens than Grok 4.6 on ARC-AGI-2

    AIGrok 4.7 used more reasoning tokens on average than Grok 4.6 on ARC-AGI-2 semi-private tasks at medium, high, and xhigh reasoning levels, raising its cost per task. Per test-pair attempt, medium used 136% more tokens, high 125% more, and xhigh 173% more, while low used 27% fewer. A chart compares the two models at xhigh on the 20 public tasks where Grok 4.7 increased token use the most.

    Image from @arcprize's post
  18. ARC PrizeOfficialAI score38

    Grok 4.7 scores 1.8% on ARC-AGI-3, trails Grok 4.6 on ARC-AGI-2

    AISpaceXAI's Grok 4.7 scored 90.2% on ARC-AGI-1 at $0.64 per task, higher than Grok 4.6, according to ARC Prize. It reached 61.4% on ARC-AGI-2 at $2.01 per task and 1.8% on ARC-AGI-3 under the standard harness ($2.7k), or 10.0% with the provider adapter harness ($4.8k), both lower than Grok 4.6 on those two benchmarks.

    Image from @arcprize's post
  19. Microsoft ResearchOfficialAI score36

    Jennifer Neville on learning from surprising AI failures and evaluation beyond benchmarks

    AIMicrosoft Research podcast host Chad Atalla interviews Jennifer Neville, a partner research manager at Microsoft, about her path into AI and her work on how evaluation exposes surprising failures in models tested beyond traditional benchmarks. The conversation also covers practical guidance for working with current AI systems and why examining underlying data matters when results defy expectations.

  20. GoogleOfficialAI score43

    EmbeddingGemma 2 Delivers Best-in-Class Performance at 740M Parameters

    AIGoogle's EmbeddingGemma 2 is a 740M-parameter embedding model that outperforms some models more than twice its size while using about 191MB to 567MB of active RAM. It offers an 8K context window, 4x larger than the first generation, and can process up to 5.5 minutes of audio, 29 images, or 58 video frames in one pass.

    Image from @Google's post
  21. Google DeepMind · The KeywordOfficialAI score72

    Google releases EmbeddingGemma 2, an open multimodal embedding model for on-device use

    AIGoogle DeepMind has released EmbeddingGemma 2, a 740-million-parameter embedding model that maps text, images, audio, and video into a shared space and runs on local hardware under an Apache 2.0 license. Matryoshka Representation Learning lets developers truncate output vectors from 768 dimensions to 512, 256, or 128, and the model supports an 8K-token context window. The model weights are available on Hugging Face and Kaggle, with Gemini Enterprise Agent Platform availability coming soon.

    Why it matters: The release shows how a 740M-parameter multimodal embedder runs locally with a 768-to-128 dimension truncation option, useful for judging on-device retrieval designs.

  22. ARC PrizeOfficialAI score14

    ARC Prize hosts Frontier AI benchmarking dinner with Snorkel AI during SF Tech Week

    AIARC Prize is hosting its first Frontier AI Benchmarking Ecosystem Dinner with Snorkel AI during SF Tech Week. The event will bring together 40 experts from nonprofit benchmarking organizations, industry, academia, and government to discuss AI benchmark design and building a more scalable, transparent evaluation ecosystem. Attendees named in the post include representatives from METR, Epoch, Vals, Artificial Analysis, Harvard, and Stanford.