How the score works
Different evaluations measure in different units at different difficulties. The leaderboard turns every result into a reading on one ability scale, so a hard evaluation and an easy one speak the same language, and a result a model never took neither adds nor subtracts.
The reference group's average is 50 and each 15 points is about one standard deviation. A score is not a percentage of questions answered. Every board shows at most 30 models.
From a result to a score
- 01
Only results an independent evaluator ran itself
A result counts only when the evaluator ran the model itself, the same way it runs every other model. Numbers a lab reports about its own model don't count, nor do runs inside a vendor's own coding agent, nor indices stitched together from other people's results. Each model is represented by one run at a reasoning setting chosen in advance, never the best of several.
- 02
Every evaluation becomes a ruler
Evaluations differ in units and difficulty, so their numbers can't simply be added. Each one gets two numbers, a difficulty and a slope, worked out from every model that took it. An evaluation is calibrated only once at least 20 models have taken it. A hard evaluation speaks loudest at the top; one that nearly every model aces stops telling models apart and matters less on its own.
- 03
Results become readings on one scale
Each result is read off its ruler as an ability reading, and every reading sits on the same scale, so different evaluations can be compared. A result above 95% only says "at least this strong" and one below 5% only "at most this strong". Evaluations of the same kind share a single vote, and a reading far from the model's other readings has its pull capped.
- 04
Domains first, then weights
A domain's score is the average of its readings, nudged slightly toward the model's overall level. A domain with no results takes the overall level, so missing evidence is neither a bonus nor a penalty. The overall score adds up the domains with the weights below.
Five domains, coding weighs most
The weights are a choice we make in the open, not a law of nature. Two evaluators re-running the same suite, or two tiers of one test, count as one vote. Vision, writing and multilingual evaluations are shown for reference and not scored yet.
Coding
40%SciCode (Artificial Analysis) · Terminal-Bench 4.0 (AA run) (Artificial Analysis) · Vibe Code Bench (Vals AI) · Code Migration (Vals AI) · LiveBench Coding (LiveBench) · ProgramBench (Vals AI) · APEX-SWE (Mercor) · Terminal-Bench 4.0 (Vals AI) · CyberBench Patch (Vals AI) · Vibe Code Bench 1–100 (Vals AI)
Research and reasoning
20%Humanity's Last Exam (AA run) (Artificial Analysis) · Chess Puzzles (Epoch AI) · FrontierMath Tiers 1–3 (Epoch AI) · Mystery Game Puzzles (Epoch AI) · LiveBench Reasoning (LiveBench) · FrontierMath Tier 4 (Epoch AI) · ProofBench (Vals AI) · Terminal-Bench Science (Vals AI) · MysteryMechanism (Vals AI)
Professional work
20%τ²-Bench Telecom (AA run) (Artificial Analysis) · τ-Bench Banking (AA run) (Artificial Analysis) · MedCode (Vals AI) · Finance Agent (Vals AI) · Harvey Legal Agent Benchmark (Vals AI) · Legal Research Bench (Vals AI) · EMB (Vals AI) · Tax Agent Bench (Vals AI) · LiveBench Data Analysis (LiveBench) · APEX-Agents (Mercor)
Knowledge and accuracy
15%AA-LCR (Artificial Analysis) · IFBench (AA run) (Artificial Analysis) · BullshitBench (BullshitBench) · SimpleQA Verified (Epoch AI) · LiveBench Language (LiveBench) · LiveBench Instruction Following (LiveBench)
Human preference
5%Arena Text (LMArena)
Enough evidence first
A rank goes only to a model whose evidence is wide enough: results from at least 2 evaluators, at least 3 kinds of coding evaluation, and at least 2 each in research, professional work and knowledge. Other models with a score are listed as "Not ranked yet" with the reason, once they have at least 5 scored results. Models released more than 18 months ago leave the current board; their pages stay. Ranks follow the score with no ties: an exact tie goes to the model with more evidence, then the newer one.
You may also wonder
QWhy not use the numbers labs publish about their own models?
Labs choose their own settings, prompts and harnesses, and they publish the runs that went well. Results only compare fairly when one evaluator runs every model the same way, so self-reported numbers never count here.
QWhat does a score of 65 mean?
It is about one standard deviation above the reference group's average, which sits at 50. The reference group is the set of models ranked on the current board. A score is a position on an ability scale, not the share of questions a model got right.
QWhy can a modest result on a hard evaluation read higher than a great one on an easy evaluation?
Because the ruler knows how hard each evaluation is. Solving a fifth of a test that almost no model can touch says more than solving nine tenths of one that most models pass. An evaluation that nearly every model aces can only say "at least this strong".
QHow big a gap is a real difference?
Treat a gap of a point or two as close to a tie that the ranking had to break: every result carries sampling noise. A wider gap backed by many results from several evaluators is much more trustworthy. Each model page shows the results behind its score, so you can see how much evidence there is.
QWhy does a model have a score but no rank?
Its evidence is too narrow: one evaluator, or too few kinds of evaluation in a domain. With little evidence, one unusual result can move a score a long way, so the model waits under "Not ranked yet" until more evaluations include it. The reason is shown next to it.
QHow many models must take an evaluation before it counts?
At least 20. With fewer, the ruler's difficulty and slope can't be pinned down, so the evaluation is listed as waiting for models and its results are shown without counting.
QWhy did a score change when nothing new was measured?
Two reasons. Every refresh fits the rulers again with all current results, so a new model on an evaluation shifts that evaluation's ruler slightly. And the scale is anchored to the models ranked today: as stronger models arrive, the same ability sits lower relative to the group.
QWhich reasoning setting represents a model?
The one closest to a fixed preference order, decided before looking at results: high, then extra high, then maximum, then a plain thinking mode, then medium, the default, low, minimal and none. When an evaluator ran several settings, the first one in that order is used, whatever it scored.
QWhat happens when an evaluator is down or slow to add new models?
Its previous results stay in place and the evaluations page marks it as unavailable. A new model that some evaluators have not run yet simply has fewer results; missing results neither help nor hurt, and the eligibility rules decide when there is enough for a rank.
QDo price or speed affect the rank?
No. Prices are shown to help you choose, from OpenRouter's public list in US dollars per million tokens, and never enter the score.
QWhen does a new model show up?
After the next daily refresh in which an evaluator has published its results. It gets a page as soon as it has a calibrated result, a place under "Not ranked yet" once it has 5, and a rank once its evidence is wide enough.
For those who want the math
Results as fractions. Percentages and fractions are used as published. Arena ratings become the expected win rate against a model at the evaluation's median rating, 1 / (1 + 10^(−(r − median) / 400)).
The rulers. Each evaluation follows a two-parameter logistic curve, p = σ(a·(θ − d)), where θ is a model's ability, d the evaluation's difficulty and a its slope. All curves and abilities are fitted together by alternating weighted least squares on the logit scale: rulers given the abilities, then abilities given the rulers, until nothing moves. Each result is weighted 4p(1 − p), the inverse of the sampling variance of logit(p), which keeps the fit close to a binomial likelihood. Results above 95% or below 5% are censored: they cost the fit something only when the curve predicts the wrong side of the bound. Weak priors keep a model with few results finite (a slight pull of its ability toward the middle) and pull slopes toward 1, within 0.2 to 6. After each round the abilities are standardized over the models with three or more results.
Readings. A result p on a ruler reads as θ = d + logit(p) / a; past the bounds it reads as the bound itself. A lower bound below the model's median exact reading counts at that median, and an upper bound above it likewise: a bound only says which side the model is on.
Domains. Readings of the same kind are averaged first. A kind more than 1.25 standard deviations of the calibration scale from the median of the model's other kinds is pulled back to that distance. A domain's score is (sum of its kinds + 0.5 × overall level) / (number of kinds + 0.5), where the overall level is the mean of all the model's kinds. The overall score weighs the five domains as above.
The scale. The ranked models form the reference group (all current models with a score while fewer than five are ranked). Each value x is printed as 50 + 15 × (x − mean) / standard deviation of the group's overall values.
- Method version
- 2026.10-irt-1
- Latest computation
- Oct 8, 11:55 PM ET