Skip to content

Evaluations we read

What each evaluation measures, who runs it, and whether it counts toward the score. Only results an evaluator ran itself, the same way for every model, are scored. How it's scored

36
scored
7
watching
63
evaluations read

Updated Oct 8, 21:18 ET

Coding

18 evaluations

Engineering work in a terminal, building and extending apps, migrating code, patching security flaws.

  • SciCodeScored

    Turns physics, chemistry and biology problems into working scientific code.

    Artificial Analysis148 modelsUpdated Oct 9, 2026

  • The same terminal tasks, re-run in Artificial Analysis's own harness; shares one vote with Vals's run.

    Artificial Analysis141 modelsUpdated Oct 9, 2026

  • Builds a web app from a written spec; the app's features are tested.

    Vals AI100 modelsUpdated Oct 7, 2026

  • Moves a working program from one language or framework to another.

    Vals AI74 modelsUpdated Oct 8, 2026

  • Fresh code-generation, completion and agentic repository tasks, averaged.

    LiveBench65 modelsUpdated Oct 7, 2026

  • Rebuilds a program from its executable alone, matching its behavior.

    Vals AI61 modelsUpdated Oct 8, 2026

  • APEX-SWEScored

    Ships realistic software engineering work end to end, pass@1.

    Mercor54 modelsUpdated Oct 9, 2026

  • Finishes whole engineering tasks alone in a command-line sandbox.

    Vals AI44 modelsUpdated Oct 7, 2026

  • Patches real security flaws in open-source projects.

    Vals AI42 modelsUpdated Oct 7, 2026

  • Keeps extending an existing app without breaking what already works.

    Vals AI22 modelsUpdated Oct 8, 2026

  • SkillsBenchWatching

    Agent tasks with and without reusable skills; run conditions still being checked.

    Vals AI34 modelsUpdated Sep 27, 2026

  • SRE BenchWatching

    Site-reliability incidents in live systems; too new to score.

    Vals AI32 modelsUpdated Oct 8, 2026

  • MirrorCodeWatching

    Reproduces software over long sessions; too few models so far.

    Epoch AI9 modelsUpdated Sep 22, 2026

  • Artificial Analysis's run of the contest problems.

    Artificial Analysis277 modelsUpdated Oct 9, 2026

  • LiveCodeBenchReference only

    Contest-style coding problems; most frontier models now solve nearly all.

    Vals AI130 modelsUpdated Sep 1, 2026

  • SWE-bench VerifiedReference only

    Fixes GitHub issues; widely trained on, so shown for reference.

    Vals AI84 modelsUpdated Sep 1, 2026

  • IOIReference only

    Olympiad algorithm problems; shown for reference.

    Vals AI41 modelsUpdated Oct 7, 2026

  • SWE-bench VerifiedReference only

    Epoch's own run of the same issue-fixing set, for cross-checking.

    Epoch AI31 modelsUpdated Jun 25, 2026

Research and reasoning

12 evaluations

Research math, logic and game puzzles, proofs, experiments and scientific workflows.

  • Expert questions across fields, answered without tools.

    Artificial Analysis467 modelsUpdated Oct 9, 2026

  • Finds the winning move from a position written out as text.

    Epoch AI141 modelsUpdated Sep 29, 2026

  • Unpublished research-level math problems with checkable answers.

    Epoch AI83 modelsUpdated Sep 29, 2026

  • Works out the right move in games whose rules it has to infer.

    Epoch AI73 modelsUpdated Sep 29, 2026

  • Fresh logic puzzles and competition math, averaged.

    LiveBench65 modelsUpdated Oct 7, 2026

  • The hardest FrontierMath tier; shares one vote with Tiers 1–3.

    Epoch AI63 modelsUpdated Sep 29, 2026

  • Writes complete mathematical proofs that are graded line by line.

    Vals AI49 modelsUpdated Oct 7, 2026

  • Runs scientific workflows and analyses in a terminal.

    Vals AI38 modelsUpdated Oct 8, 2026

  • Designs its own experiments to uncover a hidden mechanism.

    Vals AI25 modelsUpdated Oct 7, 2026

  • Biology research puzzles; still checking how it is graded.

    Vals AI25 modelsUpdated Oct 8, 2026

  • AIME (AA run)Reference only

    Competition math; saturated at the top.

    Artificial Analysis207 modelsUpdated Oct 9, 2026

  • OTIS Mock AIMEReference only

    Competition math that frontier models have nearly saturated.

    Epoch AI192 modelsUpdated Sep 29, 2026

Professional work

16 evaluations

Finance, law, tax, medical coding, spreadsheets and long agent assignments.

  • Resolves telecom support cases with tools and policies, alongside a simulated customer.

    Artificial Analysis322 modelsUpdated Oct 9, 2026

  • Handles bank customer service: policies, products and account changes; shares one vote with the telecom set.

    Artificial Analysis154 modelsUpdated Oct 9, 2026

  • MedCodeScored

    Assigns diagnosis codes from hospital records.

    Vals AI96 modelsUpdated Oct 8, 2026

  • An analyst's daily work: reading filings, adjusting numbers, building models.

    Vals AI75 modelsUpdated Oct 7, 2026

  • Legal work delivered as documents, spreadsheets and slides a lawyer can review.

    Vals AI75 modelsUpdated Oct 7, 2026

  • Finds and applies the law across practice areas.

    Vals AI74 modelsUpdated Oct 7, 2026

  • EMBScored

    Builds and repairs financial models in spreadsheets: DCF, LBO, M&A.

    Vals AI71 modelsUpdated Oct 7, 2026

  • Answers professional tax questions from facts and current rules.

    Vals AI66 modelsUpdated Oct 8, 2026

  • Reshapes tables, predicts joins and reads event sequences.

    LiveBench65 modelsUpdated Oct 7, 2026

  • Long banking, consulting and legal assignments done as an agent, pass@1.

    Mercor52 modelsUpdated Oct 9, 2026

  • MedScribeWatching

    Writes clinical notes; checking how the notes are graded.

    Vals AI98 modelsUpdated Oct 8, 2026

  • Accounting assignments; pass rates are still near the floor.

    Mercor30 modelsUpdated Oct 9, 2026

  • EBR-benchWatching

    Learns new rules during a long task and keeps using them.

    Epoch AI24 modelsUpdated Sep 29, 2026

  • LegalBenchReference only

    Short legal-reasoning tasks; close to saturated.

    Vals AI139 modelsUpdated Oct 1, 2026

  • TaxEvalReference only

    Tax questions without tools; close to saturated.

    Vals AI133 modelsUpdated Sep 1, 2026

  • CorpFinReference only

    Questions over long credit agreements; older and close to saturated.

    Vals AI122 modelsUpdated Aug 12, 2026

Knowledge and accuracy

9 evaluations

Factual recall, instruction following, long documents and spotting a false premise.

  • AA-LCRScored

    Reasons over very long documents.

    Artificial Analysis388 modelsUpdated Oct 9, 2026

  • Follows new, precisely checkable output instructions.

    Artificial Analysis330 modelsUpdated Oct 9, 2026

  • Notices a false premise instead of answering along with it.

    BullshitBench115 modelsUpdated Sep 25, 2026

  • Short factual questions answered from memory.

    Epoch AI78 modelsUpdated Sep 29, 2026

  • Word puzzles, typo fixing and plot reconstruction.

    LiveBench65 modelsUpdated Oct 7, 2026

  • Rewrites and summaries under exact formatting instructions.

    LiveBench65 modelsUpdated Oct 7, 2026

  • GPQA DiamondReference only

    Graduate-level science questions; saturated at the top.

    Epoch AI214 modelsUpdated Sep 29, 2026

  • GPQA DiamondReference only

    Vals's run of the same science questions.

    Vals AI125 modelsUpdated Sep 1, 2026

  • MMLU-ProReference only

    Broad multiple-choice knowledge; saturated at the top.

    Vals AI125 modelsUpdated Sep 1, 2026

Human preference

1 evaluation

People comparing two anonymous answers side by side.

  • People compare two anonymous answers and pick the one they prefer (style-controlled).

    LMArena385 modelsUpdated Oct 2, 2026

Composite indices

3 evaluations

Each evaluator's own summary of its evaluations: a cross-check, never scored.

  • AA Intelligence IndexReference only

    Artificial Analysis's composite of its own evaluations; a cross-check, not scored.

    Artificial Analysis492 modelsUpdated Oct 9, 2026

  • LiveBench averageReference only

    LiveBench's average across all its categories; a cross-check.

    LiveBench65 modelsUpdated Oct 7, 2026

  • Vals IndexReference only

    Vals's composite of its own benchmarks; a cross-check.

    Vals AI44 modelsUpdated Oct 7, 2026

Vision

2 evaluations

Understanding images; shown for reference, not scored yet.

  • Arena VisionReference only

    People compare answers about images.

    LMArena145 modelsUpdated Oct 2, 2026

  • MMMU ProReference only

    College-level questions about charts, diagrams and photos.

    Vals AI83 modelsUpdated Sep 1, 2026

Writing and design

1 evaluation

Creative work and design; shown for reference, not scored yet.

  • Arena WebDevReference only

    People pick the better of two generated web apps.

    LMArena124 modelsUpdated Oct 6, 2026

Multilingual

1 evaluation

Other languages; shown for reference, not scored yet.

  • MGSMReference only

    Grade-school math in ten languages.

    Vals AI64 modelsUpdated Jan 9, 2026

Evaluators

Read once a day, politely: one request at a time per site, cached, under our own crawler name. When an evaluator cannot be read, its last results stay and it is marked here.

  • Results published openly with the benchmark (code under Apache 2.0).

    Last read Oct 9, 2026

  • Benchmarking Hub data under CC BY 4.0; credited here.

    Last read Oct 9, 2026

  • Leaderboard dataset on Hugging Face under CC BY 4.0.

    Last read Oct 9, 2026

  • Public benchmark pages; results credited and linked.

    Last read Oct 9, 2026

  • MercorRead

    Public APEX leaderboards; results credited and linked.

    Last read Oct 9, 2026

  • Published in its GitHub repository under the MIT license.

    Last read Oct 9, 2026

  • Free data API with attribution; read only when ARTIFICIAL_ANALYSIS_API_KEY is set.

    Last read Oct 9, 2026

Model names, release dates, prices, context windows and open-weight flags come from OpenRouter's public model list.