Every major AI model, ranked in the open.
117 models from 18 labs on 30 third-party benchmarks with enough results to compare. Every score is measured against the best result its publisher has published on that test and averaged into one OpenCharts Index with a letter grade and an honest error band. Every number links to the publisher that produced it and says who ran it, and the models Theo runs on are marked.
36 ranked · snapshot September 8, 2026 · free, no sign-in · llms.txt · JSON
#1 by OpenCharts Index
GPT-6 Astra
OpenAI
99.0
1.0% behind the frontier on average · 14 comparable tests · 3 categories · 99.0 to 99.7 if any one test is dropped
Leads the hard set
Claude Fable 5.1
Anthropic · #2 overall
4 of 8 hardest tests · share of the frontier
Mean share of the frontier on the 8 hardest, most general tests in the set. A distance to today's frontier, not a measure of intelligence.
Waiting on more results
- GPT-6 Astra100.0 on 3 of 8
- Claude Fable 593.1 on 3 of 8
- GPT-5.6 Sol92.5 on 3 of 8
The score needs 4 of the 8 hard-set tests; these enter it as publishers add results.
Category leaders
Best at each kind of work
The third-party-published model with the highest category mean, among ranked models that sat at least half of that category's comparable tests. A category that is a single benchmark is marked: it sits beside the Index, not inside it. Tap a card to see the whole category sorted.
Reasoning
GPT-6 AstraOpenAI99.6
category mean · 7 tests in the set
Hard, novel problems: graduate science, abstract puzzles, expert exams.
Sort the table by reasoningKnowledge
GPT-6 AstraOpenAI100.0
one test in the set · shown beside the Index, not averaged into it
Recall of facts without hallucinating — short-answer accuracy.
Sort the table by knowledgeCoding
GPT-6 AstraOpenAI97.7
category mean · 7 tests in the set
Writing, fixing and shipping real code, judged by tests or users.
Sort the table by codingMath
GPT-6 AstraOpenAI99.8
category mean · 5 tests in the set
Competition and research mathematics, graded on the final answer or the proof.
Sort the table by mathAgentic
Claude Opus 4.6Anthropic91.1
category mean · 8 tests in the set
Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs.
Sort the table by agenticMultimodal
Claude Fable 5Anthropic100.0
one test in the set · shown beside the Index, not averaged into it
Understanding images alongside text, judged by human preference.
Sort the table by multimodalHuman preference
Claude Fable 5Anthropic100.0
one test in the set · shown beside the Index, not averaged into it
Blind head-to-head votes from real users on real prompts.
Sort the table by human preferenceLong context
GPT-5OpenAI96.9
one test in the set · shown beside the Index, not averaged into it
Keeping track of a story across a very long document.
Sort the table by long contextThe benchmarks
31 third-party tests, one Index
Every evaluation the Index reads, with who publishes it, who produced the numbers and what it measures. Tests in the hard set feed the hard-set score. A test with fewer than 3 listed models is shown but not counted. Open one to see every model ranked on that test alone.
Reasoning
Epoch AI (internal runs) · Epoch-run · 78 models scored · n = 198
198 PhD-level four-option questions in physics, chemistry and biology, written to be Google-proof. Epoch AI runs most models itself 16 times; random guessing scores 25%.
Frontier models cluster in the eighties and nineties; Epoch's standard errors on this test are 2 to 3 points.
Ranking on this testCenter for AI Safety & Scale AI · Leaderboard · 18 models scored · n = 2,500
2,500 expert-written questions across a hundred fields at the frontier of human knowledge. Accuracy on the full exam, from the exam's own leaderboard.
Ranking on this testARC Prize Foundation · Leaderboard · 51 models scored · n = 120
Abstract visual reasoning puzzles that are easy for people and designed to resist memorization — the benchmark built to measure general fluid intelligence. Semi-private evaluation set, verified by ARC Prize.
Scores depend on the compute budget a lab chose; ARC Prize publishes cost per task beside every score and this ranking does not.
Ranking on this testARC Prize Foundation · Leaderboard · 51 models scored · n = 100
The original Abstraction and Reasoning Corpus — grid puzzles solved from a handful of examples. Semi-private evaluation set, verified by ARC Prize.
Close to saturated at the frontier; scores depend on the compute budget a lab chose.
Ranking on this testLiveBench · Leaderboard · 1 models scored · not counted
A contamination-resistant benchmark refreshed monthly — reasoning, coding, math, data analysis, language and instruction following. Global average.
Epoch's copy of this leaderboard lags the live site, so few current models have a score here.
Ranking on this testSimpleBench / LM Council · Leaderboard · 46 models scored
Trick questions about the everyday world where humans score in the eighties — spatio-temporal reasoning, social intelligence and adversarial phrasing.
Ranking on this testCritPt / Artificial Analysis · Leaderboard · 78 models scored
Unpublished research-level physics problems graded by an official server — accuracy on the full set, as run by Artificial Analysis.
Ranking on this testCoding
Epoch AI (internal runs) · Epoch-run · 26 models scored · n = 500
500 human-validated GitHub issues the model must resolve in the real repository so the hidden tests pass. Epoch AI runs every model in one standardized scaffold.
One shared scaffold for every model; labs' own numbers with bespoke scaffolds can run 10 points higher and are not used here.
Ranking on this testAider · Leaderboard · 5 models scored · n = 225
225 of the hardest Exercism exercises across six languages, edited in place and checked by the tests.
Ranking on this testSciCode / Artificial Analysis · Leaderboard · 73 models scored · n = 338
Research-grade scientific programming: 338 sub-problems from 80 research problems across physics, chemistry, biology, materials and math that must pass unit tests. As run by Artificial Analysis.
Ranking on this testHåvard Ihle · Leaderboard · 69 models scored
Unusual machine-learning tasks the model must solve end to end by writing and running training code.
Ranking on this testCognition · Leaderboard · 27 models scored
Hard, realistic software tasks run through a coding-agent harness. Main score.
Published by Cognition, a coding-agent vendor, from its own harness; not independently reproduced.
Ranking on this testCursor · Leaderboard · 21 models scored
Real coding-agent requests drawn from everyday editor sessions, scored against the change the developer actually shipped.
Published by Cursor from its own product traffic and harness; not independently reproduced.
Ranking on this testArena · Leaderboard · 86 models scored
Blind head-to-head votes on web apps two models built from the same prompt. Arena score.
Ranking on this testMath
Epoch AI (internal runs) · Epoch-run · 78 models scored · n = 45
Competition mathematics from the OTIS Mock AIME sets — 45 problems, exact-answer graded, run 16 times by Epoch AI.
45 problems: one problem is 2.2 points, so gaps of a few points are within noise.
Ranking on this testEpoch AI (internal runs) · Epoch-run · 6 models scored
The hardest tier of the MATH competition dataset. Saturated at the frontier; useful for the long tail.
Epoch flags likely training contamination on this set; treat small gaps as meaningless.
Ranking on this testEpoch AI (internal runs) · Epoch-run · 61 models scored · n = 290
290 unpublished research-level math problems written by professional mathematicians, tiers 1–3. Epoch AI runs it.
Commissioned by OpenAI, which has access to much of the problem set; Epoch discloses this and keeps a holdout it does not share.
Ranking on this testEpoch AI (internal runs) · Epoch-run · 51 models scored · n = 48
The hardest FrontierMath tier — 48 problems that take expert mathematicians days. Epoch AI runs it.
48 problems: one problem is about 2 points. Same OpenAI commissioning disclosure as tiers 1 to 3.
Ranking on this testVals AI · Leaderboard · 61 models scored
Full written proofs, not final answers, graded for rigor.
Ranking on this testAgentic
Terminal-Bench · Leaderboard · 40 models scored
Real tasks done in a terminal — set up servers, fix builds, wrangle data — checked by tests. Best published agent harness per model.
Harnesses differ by model, so this compares model-plus-harness systems, not models alone.
Ranking on this testOpenAI (external evaluations) · Leaderboard · 8 models scored · n = 220
Economically valuable work across 44 occupations, judged by industry professionals against deliverables from their peers. Win rate on the 220-task gold set.
Authored and graded by OpenAI, which also competes on it. Treat as a lab-run leaderboard.
Ranking on this testMETR · Leaderboard · 14 models scored
The length of software task (in human-expert minutes) a model completes with 50% reliability. Shown in minutes; normalized on a log scale from one minute to the frontier.
METR publishes wide confidence intervals around each horizon; the point estimate is used here.
Ranking on this testXLANG Lab · Leaderboard · 9 models scored · n = 108
108 long-horizon computer-use workflows across real desktop and web applications; a skilled person takes a median of about 1.6 hours per task. Share of tasks completed in full at the 500-step budget.
A separate, far harder benchmark from OSWorld-Verified: binary completion is low for every model, and the leaderboard also reports a partial-credit score.
Ranking on this testMercor · Leaderboard · 49 models scored
Professional-services tasks — consulting, law, finance — completed as an agent and graded against expert rubrics. Pass@1.
Ranking on this testAndon Labs · Leaderboard · 53 models scored
Run a simulated vending business for a year from a $500 float — ordering, pricing, cash flow. Mean final balance over five runs; normalized on a log scale from the starting balance to the frontier.
Andon Labs estimates a strong human operator at roughly $63,000, so every model is far from the ceiling.
Ranking on this testStanford · Self-reported · 7 models scored · n = 40
40 professional capture-the-flag cybersecurity challenges solved unguided. Share solved.
Developer-reported: the numbers come from the labs' own model cards, collected by Epoch AI, not from an independent run.
Ranking on this testDeepResearch Bench · Leaderboard · 19 models scored · n = 100
PhD-level research briefs written by an agent with web access, graded on coverage, insight and citations.
Ranking on this testMethodology
An average you can audit
OpenCharts publishes no scores of its own for the ranked models. It reads what the publishers publish, measures every result against the best one on that test, averages, and shows its work. Coverage is printed next to every Index so you can see how much a number rests on.
- 1
Best published run
Publishers list several runs of one model on one test (effort levels, thinking budgets, agent harnesses). The best published value counts, the same rule Epoch AI's capability index applies.
- 2
Measure against the frontier
The frontier on each test is the best result its publisher lists, tracked model or not. Percent scores become the share of that frontier above random guessing, so the best model reads 100 and sitting a harder exam never lowers anyone. Arena ratings become the expected win rate against the board leader; open-ended scores (minutes, dollars) sit on a log scale from a fixed floor. A test with fewer than 3 listed models is shown but not counted.
- 3
Average inside each category
Reasoning, coding, math and agentic each get one mean. A category needs at least 2 comparable tests to join the Index; knowledge, multimodal, human preference and long context are one test each today, so they are shown beside it. A category counts for a model once it has results on at least half of that category's tests.
- 4
The OpenCharts Index
The mean of the counting category means. Balanced by category, so a lab cannot climb by stacking many tests of one kind, and no single easy test can stand in for a category.
- 5
Grade, rank, band, provenance
The grade reads the Index as a distance from the frontier. A rank needs at least 6 comparable tests across 3 counting categories, and every rank carries the range it would move within if any single test were dropped. Every score says who produced it: Epoch AI's own run, an official leaderboard, the developer itself, or an OpenCharts internal run.
Grade bands
The Index is a distance from the frontier, so every grade has a plain meaning.
- A+≥ 95Within 5% of the best published results, on average.
- A≥ 90Between 5% and 10% behind the best published results, on average.
- A-≥ 85Between 10% and 15% behind the best published results, on average.
- B+≥ 80Between 15% and 20% behind the best published results, on average.
- B≥ 75Between 20% and 25% behind the best published results, on average.
- B-≥ 70Between 25% and 30% behind the best published results, on average.
- C+≥ 65Between 30% and 35% behind the best published results, on average.
- C≥ 60Between 35% and 40% behind the best published results, on average.
- C-≥ 50Between 40% and 50% behind the best published results, on average.
- D≥ 40Between 50% and 60% behind the best published results, on average.
- F≥ 0Below 40% of the best published results, on average.
Thresholds
- Comparable test: at least 3 models scored on it. Fewer, and the score is shown but not counted.
- Counting category: the model has results on at least half of the category's comparable tests.
- Ranked: at least 6 comparable tests across 3 counting categories. Otherwise provisional.
- Category leader: the best qualifying mean among ranked, third-party-published models. Single-test categories are marked as beside the Index.
- Hard-set score: at least 4 of the 8 hardest tests (Humanity's Last Exam, ARC-AGI-2, SWE-bench Verified, FrontierMath (Tiers 1–3), FrontierMath Tier 4, Terminal-Bench, GDPval, METR Time Horizon), third-party scores only. Models with fewer are listed as waiting on more results. A distance to today's frontier on those tests, not a measure of intelligence.
- Tie: two models whose leave-one-out bands overlap, or two scores within two points in a head-to-head.
Internal scores
Where no publisher has verified an engine yet (the first-party engines Theo can run on), OpenCharts runs the same public test itself and shows the result with the dataset version, item count, repeats, grader and the engine that actually served the run. Those scores are marked on every surface, rank in the same table, and never lead a spotlight.
Questions
Questions? Answered
One number per model: how close it gets, on average, to the best published result on each third-party benchmark it has taken. Every score becomes a share of the frontier on that test (the best result reads 100), scores are averaged inside each capability category, and the Index is the mean of the categories the model has covered well enough to count. Only categories with more than one test are averaged; a category that is a single benchmark is shown beside the Index instead. Averaging by category means a model cannot climb by taking many tests of one kind, and every number links back to the publisher that produced it.
The grade is the Index read as a distance from the frontier. A+ means the model is within 5% of the best published results on average, A within 10%, A- within 15%, and so on down to F below 40%. A grey grade is provisional: the model has not covered enough comparable tests or categories to be ranked yet. Hover any grade for the exact bands and that model's coverage.
Less precise than one decimal suggests. Publishers report standard errors of two to three points on tests like GPQA Diamond, and one problem on a 45-problem exam is 2.2 points. So every Index carries a leave-one-out band: where it would land if any single test were dropped. Ranks are printed with the range of places that band covers, and two neighbours whose bands overlap are a tie, whatever order the table lists them in. Head-to-head comparisons call any gap under two points even.
From publishers other than the models' makers: Epoch AI's Benchmarking Hub, which runs several benchmarks itself and collects the rest from official leaderboards, and the Arena leaderboard dataset of blind human-preference votes. Both are published under Creative Commons Attribution. Every score says which of three things it is: run by Epoch AI, produced by an official leaderboard or third-party evaluator, or reported by the developer itself (Epoch's own classification). OpenCharts does not adjust a published number; it only normalizes and averages.
The frontier on each test is the best result the publisher lists across every model in its file, whether or not that model is listed here. A preview model or a model this site does not track can hold the frontier, and every listed model is measured against it. The frontier's value is printed with every score.
A model is ranked only once it has scores on at least six comparable benchmarks across at least three of the four Index categories, where a category counts once the model has covered at least half of its tests. Below that, an average says more about which tests were taken than about the model, so it is shown with its Index in grey and without a rank. Newly released models usually start here and move up as publishers add results.
Some engines, including the first-party TheoArca family that Theo can run on, have no third-party published scores yet. Where OpenCharts has run the same public benchmark itself, that score is shown with full run details and this label. Internal scores are ranked in the same table but never lead a spotlight, and the label never comes off until a publisher verifies the number.
The average of a model's relative scores on the hardest, most general evaluations in the set, shown once a model has published results on at least four of them. It is a way to compare frontier models on the tests built to resist memorization, and it measures distance to today's frontier on those tests, nothing more. It is not a measure of general intelligence and not a claim about any model's proximity to one.
Usually because the publishers have not run it on enough of the hardest tests yet. The score needs results on at least four of the eight hard-set benchmarks, and a model that shipped days ago often has two or three. Those models are listed as waiting on more results, with the mean they have so far, and they enter the ranking automatically once the fourth result is published.
No. Every score is expressed relative to the best published result on that test, so a model that scores 46% on an exam where 46% is the best anyone has managed reads 100 there, not 46. What does lower an Index is finishing further behind the frontier than other models on the same tests. Coverage is still printed next to every Index so you can see how much a number rests on.
The best published run counts for each model, so labs that publish many effort levels get the benefit of their best one. Agent benchmarks compare model-plus-harness systems, and harness choice can move a score by ten points. Some benchmarks are authored or graded by a lab that also competes on them, and every test page says so. Publishers' confidence intervals are not propagated; the leave-one-out band is a stand-in for them. The frontier moves whenever a better result is published, so Index values are comparable within a snapshot, not across snapshots.
The snapshot is refreshed from the publishers on a weekly schedule and whenever a notable model ships. The capture date is printed at the top of the page, on every model page and in the machine-readable files.
Yes. The same data is published as plain text at /benchmarks/llms.txt and as JSON at /benchmarks/rankings.json. Credit Epoch AI and Arena for the underlying scores as their licenses require, and OpenCharts for the Index.

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.