Humanity's Last Exam
2,500 expert-written questions across a hundred fields at the frontier of human knowledge. Accuracy on the full exam, from the exam's own leaderboard.
Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.
Publisher: Center for AI Safety & Scale AIWhat a model is asked to do
Answer an expert-written question from the frontier of a hundred academic fields, with an exact answer and a stated confidence.
For exampleGive the number of non-isomorphic groups of order 96 whose center has odd order, and state how confident you are.
Why it matters. The hardest general exam in the set. It separates models that reason from models that remember.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 18
- Who produced the numbers
- Official leaderboard
- Items graded
- n = 2,500 (0.0% each)
- Best published result
- 46.5%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (46.5%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
18 models on Humanity's Last Exam
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.0 points, so read gaps smaller than that as noise.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | Claude Fable 5.1Anthropic | 46.5% | 100.0 | xhigh | Leaderboard |
| 2 | Gemini 3.1 ProGoogle | 46.4% | 99.9 | preview | Leaderboard |
| 3 | GPT-5.4 ProOpenAI | 44.3% | 95.3 | 2026-03-05 | Leaderboard |
| 4 | Muse SparkMeta | 40.6% | 87.2 | — | Leaderboard |
| 5 | Gemini 3 ProGoogle | 37.5% | 80.7 | preview | Leaderboard |
| 6 | GPT-5.4OpenAI | 36.2% | 77.9 | 2026-03-05 · xhigh | Leaderboard |
| 7 | Claude Opus 4.7Anthropic | 36.2% | 77.8 | — | Leaderboard |
| 8 | Claude Opus 4.6Anthropic | 34.4% | 74.1 | max | Leaderboard |
| 9 | GPT-5 ProOpenAI | 31.6% | 68.0 | 2025-10-06 | Leaderboard |
| 10 | GPT-5.2OpenAI | 27.8% | 59.8 | 2025-12-11 | Leaderboard |
| 11 | GPT-5OpenAI | 25.3% | 54.5 | 2025-08-07 · high | Leaderboard |
| 12 | Claude Opus 4.5Anthropic | 25.2% | 54.2 | 20251101 | Leaderboard |
| 13 | Kimi K2.5Moonshot AI | 24.4% | 52.4 | — | Leaderboard |
| 14 | GPT-5.1OpenAI | 23.7% | 50.9 | 2025-11-13 | Leaderboard |
| 15 | GPT-5 miniOpenAI | 19.4% | 41.8 | 2025-08-07 | Leaderboard |
| 16 | Claude Sonnet 4.5Anthropic | 13.7% | 29.5 | 20250929 | Leaderboard |
| 17 | Claude Opus 4.1Anthropic | 11.5% | 24.8 | 20250805 | Leaderboard |
| 18 | Gemini 3.1 Flash-LiteGoogle | 8.6% | 18.6 | — | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.