All benchmarks
MathHard set

FrontierMath Tier 4

The hardest FrontierMath tier — 48 problems that take expert mathematicians days. Epoch AI runs it.

Competition and research mathematics, graded on the final answer or the proof. One of 5 comparable tests in math, a category the Index averages.

Read with care. 48 problems: one problem is about 2 points. Same OpenAI commissioning disclosure as tiers 1 to 3.

Publisher: Epoch AI (internal runs)

What a model is asked to do

Solve the hardest FrontierMath tier: problems that take expert mathematicians days.

For exampleCompute the exact value of an infinite series built from the coefficients of a modular form, and confirm it by a second method.

Why it matters. The deepest end of the set. Progress here signals genuinely new capability.

The example is original and illustrative, not an item from the dataset.

Models scored here
51
Who produced the numbers
Run by Epoch AI
Items graded
n = 48 (2.1% each)
Best published result
97.6%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (97.6%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.

Snapshot September 8, 2026

Ranking on this test

51 models on FrontierMath Tier 4

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 2.1 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1GPT-6 AstraOpenAI97.6%Epoch-run
2Claude Fable 5Anthropic90.2%Epoch-run
3Claude Fable 5.1Anthropic87.8%Epoch-run
4GPT-5.6 SolOpenAI82.9%Epoch-run
5GPT-5.6 Sol ProOpenAI80.5%Epoch-run
6GPT-5.5 ProOpenAI78.0%Epoch-run
7Claude Opus 5Anthropic73.2%Epoch-run
8GPT-5.5OpenAI72.5%Epoch-run
9GPT-5.6 TerraOpenAI70.7%Epoch-run
10GPT-5.6 LunaOpenAI61.0%Epoch-run
11GPT-5.4 ProOpenAI58.5%Epoch-run
12Claude Opus 4.8Anthropic56.1%Epoch-run
13GPT-5.4OpenAI49.0%Epoch-run
14Qwen3.8 MaxAlibaba46.3%Epoch-run
15GPT-5.2 ProOpenAI46.0%Epoch-run
16Kimi K3Moonshot AI39.0%Epoch-run
17Gemini 3.7 FlashGoogle36.6%Epoch-run
18Qwen3.7 MaxAlibaba34.1%Epoch-run
19Grok 4.6xAI31.7%Epoch-run
20Claude Opus 4.7Anthropic31.7%Epoch-run
21GPT-5.2OpenAI31.7%Epoch-run
22Claude Sonnet 5Anthropic29.3%Epoch-run
23GLM 5.2Zhipu AI29.3%Epoch-run
24GLM 5.3Zhipu AI29.3%Epoch-run
25Claude Opus 4.6Anthropic26.8%Epoch-run
26Gemini 3.1 ProGoogle26.8%Epoch-run
27Gemini 3.5 FlashGoogle26.8%Epoch-run
28DeepSeek V4 ProDeepSeek26.8%Epoch-run
29Kimi K2.6Moonshot AI25.6%Epoch-run
30Grok 4.5xAI24.4%Epoch-run
31DeepSeek V4 FlashDeepSeek24.4%Epoch-run
32Gemini 3.6 FlashGoogle22.0%Epoch-run
33GPT-5OpenAI22.0%Epoch-run
34GPT-5 ProOpenAI19.5%Epoch-run
35Gemini 3 FlashGoogle17.1%Epoch-run
36Grok 4.20xAI17.1%Epoch-run
37Inkling SmallThinking Machines17.1%Epoch-run
38GLM 5.3 FlashZhipu AI17.1%Epoch-run
39Grok 4.3xAI14.6%Epoch-run
40Kimi K2.7 CodeMoonshot AI12.2%Epoch-run
41GPT-5 miniOpenAI12.2%Epoch-run
42GPT-5.4 nanoOpenAI12.2%Epoch-run
43GPT-5.4 miniOpenAI9.8%Epoch-run
44Claude Opus 4.5Anthropic4.9%Epoch-run
45InklingThinking Machines4.9%Epoch-run
46Claude Sonnet 4.5Anthropic2.4%Epoch-run
47Claude Opus 4.1Anthropic2.4%Epoch-run
48GPT-5 nanoOpenAI2.4%Epoch-run
49GPT-5.5 InstantOpenAI2.4%Epoch-run
50Gemini 2.5 ProGoogle0.0%Epoch-run
51Gemini 3.5 Flash-LiteGoogle0.0%Epoch-run
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.