Methodology
How the OpenCharts Index works
OpenCharts runs no evaluations of its own for the ranked models. It reads what third-party publishers publish, measures every score against the frontier on its test, averages by category, and shows every step. This page is the contract: the thresholds and grade bands below are read from the same code that ranks the models, and the facts about the current snapshot are computed from the same data. Snapshot September 8, 2026.
Step 0
Sources, licenses and provenance
Third-party scores come from two publishers, both under Creative Commons Attribution 4.0, refreshed on a weekly schedule by a script that matches publisher identifiers to this registry through an explicit alias list. Nothing is fuzzy-matched; an unrecognized identifier is reported and left out. Published numbers are never adjusted; the only operations applied are the normalization and averaging described below.
- Epoch AI — AI Benchmarking Hub (CC BY 4.0): Epoch AI's own runs plus results it collects from official leaderboards and from developers' own publications, one file per benchmark, with a source link on every row. Aider Polyglot, LiveBench and Terminal-Bench subsets are Apache 2.0.
- Arena (LMArena) — Leaderboard Dataset (CC BY 4.0): the style-controlled text and vision boards and the WebDev board, from blind human-preference votes.
“Third-party published” and “independently re-run” are different claims, so every score carries its provenance, in the three-way classification Epoch AI applies to its own hub. In this snapshot:
- Run by Epoch AI357 scores on 7 benchmarks. Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.
- Official leaderboard932 scores on 23 benchmarks. Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
- Developer-reported7 scores on 1 benchmark. Published by the model's own maker in a model card or launch post and collected by Epoch AI. Not an independent run. (Cybench.)
Developer-reported numbers are shown, marked in amber on every surface, and counted like any other published score; they are never described as independent. 31 benchmarks are read today, grouped into 8 categories: Reasoning, Knowledge, Coding, Math, Agentic, Multimodal, Human preference, Long context. Every test page states its publisher, its provenance, the number of items graded and any conflict of interest a careful reader needs.
Step 1
Best published run
Publishers often list several runs of one model on one test: reasoning effort levels, thinking budgets, or different agent harnesses. The best published value counts and the variant is recorded next to the score. This is the convention Epoch AI's own capability index uses (the best score across settings for a model) and the one the Terminal-Bench leaderboard applies across harnesses. Its cost is stated under Limitations: labs that publish many settings get the benefit of their best one.
Step 2
Measured against the frontier
Every score becomes a number from 0 to 100 that means the same thing on every test: how close the model came to the best published result on that test. The frontier is the best result the publisher lists across its whole file, whether or not that model is tracked here, recorded at refresh time and printed with every score. Sitting a harder exam therefore never lowers a model; finishing further behind the frontier than other models does. In this snapshot the frontiers on 3 tests are held by a model this site does not list (LiveBench, METR Time Horizon, Fiction.LiveBench), so no listed model reads 100 there.
- Percent scores (a share of items solved) become the share of the frontier above chance: (score − chance) / (frontier − chance). This is how Epoch AI's capability index rescales chance to zero. A model at 46% on a free-answer exam whose frontier is 46% reads 100; a model at 23% reads 50. On a four-option exam where the frontier is 90%, a model at 57.5% reads 50, because its margin above guessing (32.5 points) is half the frontier's (65). Chance is applied only where guessing earns something (GPQA Diamond, 25%); free-answer, proof, code and agent tests have no chance term.
- Arena ratings (Code Arena (WebDev), Vision Arena, Text Arena) are Bradley-Terry ratings, so a gap in points has a fixed meaning: the expected win rate between the two models. Each rating becomes the expected win rate against the board leader under the Elo formula, scaled so parity reads 100: 43% for 50 points behind (reads 86), 36% for 100 behind (reads 72), 9% for 400 behind (reads 18). Unlike min-to-max scaling, this is the same on every board and does not depend on which weak model happens to be listed.
- Open-ended values (METR Time Horizon, floor 1 min; Vending-Bench 2, floor $500) are heavy-tailed, so they are placed on a log scale from a fixed, published floor to the frontier: ln(value / floor) / ln(frontier / floor). A doubling counts the same anywhere on the scale, and the floor is an anchor the publisher defines (one minute of task length; the starting balance), not the weakest listed model.
- Comparable tests only: a test with fewer than 3 listed models scored has no field to compare against, so its scores are shown on the model page but never counted.
- A lower-is-better benchmark would be inverted; none is in the set today.
Steps 3 and 4
Category means and the OpenCharts Index
A model's relative scores are averaged inside each category. A category is indexed only when the snapshot holds at least 2 comparable tests in it: a category that is a single benchmark would make that one test worth a whole category, so it is shown beside the Index and never averaged into it. In this snapshot the Index averages 4 categories (Reasoning, Coding, Math, Agentic); Knowledge, Multimodal, Human preference, Long context sit beside it.
Within an indexed category, a model's mean counts once the model has results on at least 50% of that category's comparable tests, so a single easy test can never stand in for a whole category; means below that line are still shown, faded, on the model page. The OpenCharts Index is the mean of the counting category means. Averaging by category first means a lab cannot climb by publishing many tests of one kind. A model with scores but no counting category keeps a provisional Index over what it has.
Step 5
Grades, ranking and thresholds
Because the Index is a distance from the frontier, every grade has a plain meaning. Hover any grade on the site for the band and that model's coverage.
- A+≥ 95Within 5% of the best published results, on average.
- A≥ 90Between 5% and 10% behind the best published results, on average.
- A-≥ 85Between 10% and 15% behind the best published results, on average.
- B+≥ 80Between 15% and 20% behind the best published results, on average.
- B≥ 75Between 20% and 25% behind the best published results, on average.
- B-≥ 70Between 25% and 30% behind the best published results, on average.
- C+≥ 65Between 30% and 35% behind the best published results, on average.
- C≥ 60Between 35% and 40% behind the best published results, on average.
- C-≥ 50Between 40% and 50% behind the best published results, on average.
- D≥ 40Between 50% and 60% behind the best published results, on average.
- F≥ 0Below 40% of the best published results, on average.
- Ranked: at least 6 comparable tests across 3 counting categories. Anything less is shown as provisional with its Index in grey; a provisional Index says more about which tests were taken than about the model.
- Category leader: the third-party published, ranked model with the highest qualifying mean in that category. Leaders of categories that sit beside the Index are shown and marked as such.
- Confidence: high at 12+ comparable tests across 4+ counting categories, medium at 7+ across 3+, otherwise low.
Thresholds and bands were tuned against the first live snapshots on 2026-09-08 and are versioned with the code. The first version scored tests on an absolute scale, which punished models for sitting the hardest exams; the relative scale replaced it the same day. The frontier, chance correction, Elo win rates, fixed log floors, indexed-category rule and rank bands followed after an external review of the method.
Step 6
How precise a rank is
Less precise than one decimal suggests. Publishers report standard errors of two to three points on tests like GPQA Diamond; one problem on a 45-problem exam is 2.2 points; and adjacent ranks are often separated by less than a point. So the site never shows a rank as more precise than the data supports:
- Leave-one-out band: every Index is recomputed with each of its counted tests dropped in turn. The lowest and highest results are the band, printed on the model page. A model resting on few tests moves a lot; a model with broad coverage barely moves.
- Rank band: the places a model could hold across its Index band, with every other ranked model held at its point estimate. It is printed under every rank. Two neighbours whose bands overlap are a tie, whatever order the table lists them in.
- Head to head: on a single shared test, a gap under 2 points is called even.
Publishers' own confidence intervals are not propagated (they are not published uniformly), so the band is a stand-in for them, not a substitute. The frontier moves whenever a better result is published, which means Index values are comparable within a snapshot and not across snapshots.
Score
The hard-set score
The mean relative score across the 8 hardest, most general evaluations in the set: Humanity's Last Exam, ARC-AGI-2, SWE-bench Verified, FrontierMath (Tiers 1–3), FrontierMath Tier 4, Terminal-Bench, GDPval, METR Time Horizon. It is shown once a model has published results on at least 4 of them, and it counts third-party scores only. Models with fewer results (at least 2) are listed as waiting on more results, with the mean they have so far, and enter the ranking automatically once the next result is published; a model that shipped days ago is usually there. The set was chosen editorially, for tests built to resist memorization and to span reasoning, mathematics, software, agentic work and long tasks; membership is data in the registry and every member is marked on its test page.
It measures distance to today's frontier on those tests, nothing more. It is not a measure of general intelligence and not a claim about any model's proximity to one. Coverage (how many of the 8 a model has) is always printed beside it.
Labels
Verification and provenance
Every score carries its verification and its provenance. Third-party published means a publisher other than the model's maker produced the number; its provenance says whether Epoch AI ran it, an official leaderboard or third-party evaluator produced it, or the developer reported it. Internal means OpenCharts ran the evaluation itself. A model's verification is derived from its scores: third-party when every counted score is, internal when any is. Internal models rank in the same table, carry the amber label on every surface, and never lead the hard-set or category spotlights. A developer-reported number is never described as independent anywhere on the site.
Protocol
Internal evaluation protocol
Some engines, including the first-party engines Theo can run on, have no third-party published scores yet. Their independent verification is TBD. Where OpenCharts has run the same public benchmark itself, the score is shown with this provenance on every card:
- Benchmarks: only tests with a public item set and a deterministic grader are run internally (GPQA Diamond, Humanity's Last Exam, ARC-AGI-2, ARC-AGI-1, SimpleQA Verified, Mock AIME 2024–2025, MATH Level 5). Agent-harness benchmarks stay TBD.
- Prompting: versioned zero-shot templates (currently v1): the simple-evals multiple-choice format, boxed final answers for exact-answer math, the official exam response format for open answers, JSON grids for abstraction puzzles.
- Grading: answer letter, normalized boxed answer, or exact grid match. Open answers use exact match first and an equality-checker judge second, the same convention independent publishers use.
- Served engine: the engine that actually answered is recorded per sample. A run is attributed only when the requested engine served the majority of samples; otherwise no score is written.
- Provenance: dataset version, item count, repeats, grader, run date. Pass@1 over all attempts; an ungradeable answer counts as wrong.
Internal scores never feed the hard-set score or the category leaders, and the label stays until a publisher verifies the engine. The eval runner and its graders are unit-tested and versioned with this site.
Read before citing
Limitations
These are the places where the method can mislead a reader who does not know them. They are stated here once and repeated on the surfaces they affect.
- Best-variant selection: the best published run counts, so a lab that publishes many effort levels or harness configurations gets the benefit of its best one, and a lab that publishes a single default setting does not. The variant is printed beside every score.
- Harness variance: agent benchmarks (terminals, desktops, software tasks) measure a model inside a scaffold. Harness choice can move a score by ten points, so those tests compare model-plus-harness systems, not models alone. Epoch AI's own runs use one scaffold for every model; leaderboards that accept submissions do not.
- Lab-authored benchmarks: some tests were written, commissioned or graded by a lab that also competes on them. Every such test page carries a caveat naming the conflict, and the number is still a third-party publication in the field's terms, but it is not the same evidence as a neutral run.
- Developer-reported scores: a few benchmarks are only available as the labs' own reported numbers, collected by Epoch AI. They are marked in amber wherever they appear and counted like any other published score.
- No propagated uncertainty: publishers' confidence intervals are not carried into the Index. The leave-one-out band and the rank band stand in for them, and the tie rules exist because of it.
- Frontier drift: because every score is measured against the best published result, a new record on one test lowers every other model's reading there without any of them changing. Compare Index values within a snapshot; use the raw scores to compare across time.
- Coverage gaps: publishers run new models on a few tests first, so a model's coverage says a great deal about how recently it shipped. Provisional models, the coverage counts and the “waiting on more results” list exist so a thin average is never read as a verdict.
- Editorial choices: the benchmark set, the category assignments, the hard set and every threshold on this page are choices, versioned with the code, not facts about the models.
Freshness
Update cadence
The third-party snapshot is rebuilt on a weekly schedule and whenever a notable model ships; each refresh is reviewed before it is published, and its capture date is printed on every page. Internal runs are added as they complete. Current snapshot: September 8, 2026 (1296 third-party scores, 0 internal).
For agents
Machine-readable files
The ranking is published in two formats from the same data the pages render:
- /benchmarks/llms.txt (plain text, llms.txt style)
- /benchmarks/rankings.json (JSON: every model, rank band, category mean, score, frontier, source and provenance)
Attribution
Credit where the numbers come from
- Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from https://epoch.ai/benchmarks. License: CC BY 4.0. epoch.ai/benchmarks
- Arena Leaderboard Dataset, lmarena-ai (Hugging Face, CC BY 4.0). Live boards at https://arena.ai/leaderboard. License: CC BY 4.0. the dataset
- The OpenCharts Index, grades, rank bands, category leaders, hard-set score and internal evaluations are published by OpenCharts. Reuse them with a link to this page; the underlying scores keep their publishers' licenses. Lab and model names are nominative: they identify the systems being ranked and imply no endorsement.

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.