OpenCharts Benchmarks

Every major AI model, ranked in the open.

117 models from 18 labs on 30 third-party benchmarks with enough results to compare. Every score is measured against the best result its publisher has published on that test and averaged into one OpenCharts Index with a letter grade and an honest error band. Every number links to the publisher that produced it and says who ran it, and the models Theo runs on are marked.

36 ranked · snapshot September 8, 2026 · free, no sign-in · llms.txt · JSON

Category leaders

Best at each kind of work

The third-party-published model with the highest category mean, among ranked models that sat at least half of that category's comparable tests. A category that is a single benchmark is marked: it sits beside the Index, not inside it. Tap a card to see the whole category sorted.

Reasoning

GPT-6 AstraOpenAI

99.6

category mean · 7 tests in the set

Hard, novel problems: graduate science, abstract puzzles, expert exams.

Sort the table by reasoning

Knowledge

GPT-6 AstraOpenAI

100.0

one test in the set · shown beside the Index, not averaged into it

Recall of facts without hallucinating — short-answer accuracy.

Sort the table by knowledge

Coding

GPT-6 AstraOpenAI

97.7

category mean · 7 tests in the set

Writing, fixing and shipping real code, judged by tests or users.

Sort the table by coding

Math

GPT-6 AstraOpenAI

99.8

category mean · 5 tests in the set

Competition and research mathematics, graded on the final answer or the proof.

Sort the table by math

Agentic

Claude Opus 4.6Anthropic

91.1

category mean · 8 tests in the set

Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs.

Sort the table by agentic

Multimodal

Claude Fable 5Anthropic

100.0

one test in the set · shown beside the Index, not averaged into it

Understanding images alongside text, judged by human preference.

Sort the table by multimodal

Human preference

Claude Fable 5Anthropic

100.0

one test in the set · shown beside the Index, not averaged into it

Blind head-to-head votes from real users on real prompts.

Sort the table by human preference

Long context

GPT-5OpenAI

96.9

one test in the set · shown beside the Index, not averaged into it

Keeping track of a story across a very long document.

Sort the table by long context

The benchmarks

31 third-party tests, one Index

Every evaluation the Index reads, with who publishes it, who produced the numbers and what it measures. Tests in the hard set feed the hard-set score. A test with fewer than 3 listed models is shown but not counted. Open one to see every model ranked on that test alone.

Reasoning

GPQA Diamond

Epoch AI (internal runs) · Epoch-run · 78 models scored · n = 198

198 PhD-level four-option questions in physics, chemistry and biology, written to be Google-proof. Epoch AI runs most models itself 16 times; random guessing scores 25%.

Frontier models cluster in the eighties and nineties; Epoch's standard errors on this test are 2 to 3 points.

Ranking on this test
Humanity's Last ExamHard set

Center for AI Safety & Scale AI · Leaderboard · 18 models scored · n = 2,500

2,500 expert-written questions across a hundred fields at the frontier of human knowledge. Accuracy on the full exam, from the exam's own leaderboard.

Ranking on this test
ARC-AGI-2Hard set

ARC Prize Foundation · Leaderboard · 51 models scored · n = 120

Abstract visual reasoning puzzles that are easy for people and designed to resist memorization — the benchmark built to measure general fluid intelligence. Semi-private evaluation set, verified by ARC Prize.

Scores depend on the compute budget a lab chose; ARC Prize publishes cost per task beside every score and this ranking does not.

Ranking on this test
ARC-AGI-1

ARC Prize Foundation · Leaderboard · 51 models scored · n = 100

The original Abstraction and Reasoning Corpus — grid puzzles solved from a handful of examples. Semi-private evaluation set, verified by ARC Prize.

Close to saturated at the frontier; scores depend on the compute budget a lab chose.

Ranking on this test
LiveBench

LiveBench · Leaderboard · 1 models scored · not counted

A contamination-resistant benchmark refreshed monthly — reasoning, coding, math, data analysis, language and instruction following. Global average.

Epoch's copy of this leaderboard lags the live site, so few current models have a score here.

Ranking on this test
SimpleBench

SimpleBench / LM Council · Leaderboard · 46 models scored

Trick questions about the everyday world where humans score in the eighties — spatio-temporal reasoning, social intelligence and adversarial phrasing.

Ranking on this test
CritPt

CritPt / Artificial Analysis · Leaderboard · 78 models scored

Unpublished research-level physics problems graded by an official server — accuracy on the full set, as run by Artificial Analysis.

Ranking on this test

Coding

SWE-bench VerifiedHard set

Epoch AI (internal runs) · Epoch-run · 26 models scored · n = 500

500 human-validated GitHub issues the model must resolve in the real repository so the hidden tests pass. Epoch AI runs every model in one standardized scaffold.

One shared scaffold for every model; labs' own numbers with bespoke scaffolds can run 10 points higher and are not used here.

Ranking on this test
Aider Polyglot

Aider · Leaderboard · 5 models scored · n = 225

225 of the hardest Exercism exercises across six languages, edited in place and checked by the tests.

Ranking on this test
SciCode

SciCode / Artificial Analysis · Leaderboard · 73 models scored · n = 338

Research-grade scientific programming: 338 sub-problems from 80 research problems across physics, chemistry, biology, materials and math that must pass unit tests. As run by Artificial Analysis.

Ranking on this test
WeirdML

Håvard Ihle · Leaderboard · 69 models scored

Unusual machine-learning tasks the model must solve end to end by writing and running training code.

Ranking on this test
FrontierCode

Cognition · Leaderboard · 27 models scored

Hard, realistic software tasks run through a coding-agent harness. Main score.

Published by Cognition, a coding-agent vendor, from its own harness; not independently reproduced.

Ranking on this test
CursorBench

Cursor · Leaderboard · 21 models scored

Real coding-agent requests drawn from everyday editor sessions, scored against the change the developer actually shipped.

Published by Cursor from its own product traffic and harness; not independently reproduced.

Ranking on this test
Code Arena (WebDev)

Arena · Leaderboard · 86 models scored

Blind head-to-head votes on web apps two models built from the same prompt. Arena score.

Ranking on this test

Math

Agentic

Terminal-BenchHard set

Terminal-Bench · Leaderboard · 40 models scored

Real tasks done in a terminal — set up servers, fix builds, wrangle data — checked by tests. Best published agent harness per model.

Harnesses differ by model, so this compares model-plus-harness systems, not models alone.

Ranking on this test
GDPvalHard set

OpenAI (external evaluations) · Leaderboard · 8 models scored · n = 220

Economically valuable work across 44 occupations, judged by industry professionals against deliverables from their peers. Win rate on the 220-task gold set.

Authored and graded by OpenAI, which also competes on it. Treat as a lab-run leaderboard.

Ranking on this test
METR Time HorizonHard set

METR · Leaderboard · 14 models scored

The length of software task (in human-expert minutes) a model completes with 50% reliability. Shown in minutes; normalized on a log scale from one minute to the frontier.

METR publishes wide confidence intervals around each horizon; the point estimate is used here.

Ranking on this test
OSWorld 2.0

XLANG Lab · Leaderboard · 9 models scored · n = 108

108 long-horizon computer-use workflows across real desktop and web applications; a skilled person takes a median of about 1.6 hours per task. Share of tasks completed in full at the 500-step budget.

A separate, far harder benchmark from OSWorld-Verified: binary completion is low for every model, and the leaderboard also reports a partial-credit score.

Ranking on this test
APEX-Agents

Mercor · Leaderboard · 49 models scored

Professional-services tasks — consulting, law, finance — completed as an agent and graded against expert rubrics. Pass@1.

Ranking on this test
Vending-Bench 2

Andon Labs · Leaderboard · 53 models scored

Run a simulated vending business for a year from a $500 float — ordering, pricing, cash flow. Mean final balance over five runs; normalized on a log scale from the starting balance to the frontier.

Andon Labs estimates a strong human operator at roughly $63,000, so every model is far from the ceiling.

Ranking on this test
CybenchSelf-reported

Stanford · Self-reported · 7 models scored · n = 40

40 professional capture-the-flag cybersecurity challenges solved unguided. Share solved.

Developer-reported: the numbers come from the labs' own model cards, collected by Epoch AI, not from an independent run.

Ranking on this test
DeepResearch Bench

DeepResearch Bench · Leaderboard · 19 models scored · n = 100

PhD-level research briefs written by an agent with web access, graded on coverage, insight and citations.

Ranking on this test

Methodology

An average you can audit

OpenCharts publishes no scores of its own for the ranked models. It reads what the publishers publish, measures every result against the best one on that test, averages, and shows its work. Coverage is printed next to every Index so you can see how much a number rests on.

  1. 1

    Best published run

    Publishers list several runs of one model on one test (effort levels, thinking budgets, agent harnesses). The best published value counts, the same rule Epoch AI's capability index applies.

  2. 2

    Measure against the frontier

    The frontier on each test is the best result its publisher lists, tracked model or not. Percent scores become the share of that frontier above random guessing, so the best model reads 100 and sitting a harder exam never lowers anyone. Arena ratings become the expected win rate against the board leader; open-ended scores (minutes, dollars) sit on a log scale from a fixed floor. A test with fewer than 3 listed models is shown but not counted.

  3. 3

    Average inside each category

    Reasoning, coding, math and agentic each get one mean. A category needs at least 2 comparable tests to join the Index; knowledge, multimodal, human preference and long context are one test each today, so they are shown beside it. A category counts for a model once it has results on at least half of that category's tests.

  4. 4

    The OpenCharts Index

    The mean of the counting category means. Balanced by category, so a lab cannot climb by stacking many tests of one kind, and no single easy test can stand in for a category.

  5. 5

    Grade, rank, band, provenance

    The grade reads the Index as a distance from the frontier. A rank needs at least 6 comparable tests across 3 counting categories, and every rank carries the range it would move within if any single test were dropped. Every score says who produced it: Epoch AI's own run, an official leaderboard, the developer itself, or an OpenCharts internal run.

Read the full methodology

Grade bands

The Index is a distance from the frontier, so every grade has a plain meaning.

  • A+95Within 5% of the best published results, on average.
  • A90Between 5% and 10% behind the best published results, on average.
  • A-85Between 10% and 15% behind the best published results, on average.
  • B+80Between 15% and 20% behind the best published results, on average.
  • B75Between 20% and 25% behind the best published results, on average.
  • B-70Between 25% and 30% behind the best published results, on average.
  • C+65Between 30% and 35% behind the best published results, on average.
  • C60Between 35% and 40% behind the best published results, on average.
  • C-50Between 40% and 50% behind the best published results, on average.
  • D40Between 50% and 60% behind the best published results, on average.
  • F0Below 40% of the best published results, on average.

Thresholds

  • Comparable test: at least 3 models scored on it. Fewer, and the score is shown but not counted.
  • Counting category: the model has results on at least half of the category's comparable tests.
  • Ranked: at least 6 comparable tests across 3 counting categories. Otherwise provisional.
  • Category leader: the best qualifying mean among ranked, third-party-published models. Single-test categories are marked as beside the Index.
  • Hard-set score: at least 4 of the 8 hardest tests (Humanity's Last Exam, ARC-AGI-2, SWE-bench Verified, FrontierMath (Tiers 1–3), FrontierMath Tier 4, Terminal-Bench, GDPval, METR Time Horizon), third-party scores only. Models with fewer are listed as waiting on more results. A distance to today's frontier on those tests, not a measure of intelligence.
  • Tie: two models whose leave-one-out bands overlap, or two scores within two points in a head-to-head.

Internal scores

Where no publisher has verified an engine yet (the first-party engines Theo can run on), OpenCharts runs the same public test itself and shows the result with the dataset version, item count, repeats, grader and the engine that actually served the run. Those scores are marked on every surface, rank in the same table, and never lead a spotlight.

Questions

Questions? Answered

One number per model: how close it gets, on average, to the best published result on each third-party benchmark it has taken. Every score becomes a share of the frontier on that test (the best result reads 100), scores are averaged inside each capability category, and the Index is the mean of the categories the model has covered well enough to count. Only categories with more than one test are averaged; a category that is a single benchmark is shown beside the Index instead. Averaging by category means a model cannot climb by taking many tests of one kind, and every number links back to the publisher that produced it.

The grade is the Index read as a distance from the frontier. A+ means the model is within 5% of the best published results on average, A within 10%, A- within 15%, and so on down to F below 40%. A grey grade is provisional: the model has not covered enough comparable tests or categories to be ranked yet. Hover any grade for the exact bands and that model's coverage.

Less precise than one decimal suggests. Publishers report standard errors of two to three points on tests like GPQA Diamond, and one problem on a 45-problem exam is 2.2 points. So every Index carries a leave-one-out band: where it would land if any single test were dropped. Ranks are printed with the range of places that band covers, and two neighbours whose bands overlap are a tie, whatever order the table lists them in. Head-to-head comparisons call any gap under two points even.

From publishers other than the models' makers: Epoch AI's Benchmarking Hub, which runs several benchmarks itself and collects the rest from official leaderboards, and the Arena leaderboard dataset of blind human-preference votes. Both are published under Creative Commons Attribution. Every score says which of three things it is: run by Epoch AI, produced by an official leaderboard or third-party evaluator, or reported by the developer itself (Epoch's own classification). OpenCharts does not adjust a published number; it only normalizes and averages.

The frontier on each test is the best result the publisher lists across every model in its file, whether or not that model is listed here. A preview model or a model this site does not track can hold the frontier, and every listed model is measured against it. The frontier's value is printed with every score.

A model is ranked only once it has scores on at least six comparable benchmarks across at least three of the four Index categories, where a category counts once the model has covered at least half of its tests. Below that, an average says more about which tests were taken than about the model, so it is shown with its Index in grey and without a rank. Newly released models usually start here and move up as publishers add results.

Some engines, including the first-party TheoArca family that Theo can run on, have no third-party published scores yet. Where OpenCharts has run the same public benchmark itself, that score is shown with full run details and this label. Internal scores are ranked in the same table but never lead a spotlight, and the label never comes off until a publisher verifies the number.

The average of a model's relative scores on the hardest, most general evaluations in the set, shown once a model has published results on at least four of them. It is a way to compare frontier models on the tests built to resist memorization, and it measures distance to today's frontier on those tests, nothing more. It is not a measure of general intelligence and not a claim about any model's proximity to one.

Usually because the publishers have not run it on enough of the hardest tests yet. The score needs results on at least four of the eight hard-set benchmarks, and a model that shipped days ago often has two or three. Those models are listed as waiting on more results, with the mean they have so far, and they enter the ranking automatically once the fourth result is published.

No. Every score is expressed relative to the best published result on that test, so a model that scores 46% on an exam where 46% is the best anyone has managed reads 100 there, not 46. What does lower an Index is finishing further behind the frontier than other models on the same tests. Coverage is still printed next to every Index so you can see how much a number rests on.

The best published run counts for each model, so labs that publish many effort levels get the benefit of their best one. Agent benchmarks compare model-plus-harness systems, and harness choice can move a score by ten points. Some benchmarks are authored or graded by a lab that also competes on them, and every test page says so. Publishers' confidence intervals are not propagated; the leave-one-out band is a stand-in for them. The frontier moves whenever a better result is published, so Index values are comparable within a snapshot, not across snapshots.

The snapshot is refreshed from the publishers on a weekly schedule and whenever a notable model ships. The capture date is printed at the top of the page, on every model page and in the machine-readable files.

Yes. The same data is published as plain text at /benchmarks/llms.txt and as JSON at /benchmarks/rankings.json. Credit Epoch AI and Arena for the underlying scores as their licenses require, and OpenCharts for the Index.

Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.