All benchmarks
MathInternally runnable

Mock AIME 2024–2025

Competition mathematics from the OTIS Mock AIME sets — 45 problems, exact-answer graded, run 16 times by Epoch AI.

Competition and research mathematics, graded on the final answer or the proof. One of 5 comparable tests in math, a category the Index averages.

Read with care. 45 problems: one problem is 2.2 points, so gaps of a few points are within noise.

Publisher: Epoch AI (internal runs)

What a model is asked to do

Solve competition mathematics problems that have a single exact integer answer.

For exampleFind the number of ordered pairs of positive integers (a, b) with a + b = 2026 such that a and b share no prime factor.

Why it matters. Olympiad-style problems need long, careful chains of reasoning, and there is no partial credit.

The example is original and illustrative, not an item from the dataset.

Models scored here
78
Who produced the numbers
Run by Epoch AI
Items graded
n = 45 (2.2% each)
Best published result
100.0%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (100.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.

Snapshot September 8, 2026

Ranking on this test

78 models on Mock AIME 2024–2025

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 2.2 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1GPT-6 AstraOpenAI100.0%Epoch-run
2Claude Fable 5.1Anthropic100.0%Epoch-run
3Claude Fable 5Anthropic100.0%Epoch-run
4GPT-5.6 SolOpenAI100.0%Epoch-run
5GPT-5.6 TerraOpenAI99.7%Epoch-run
6Qwen3.8 MaxAlibaba99.4%Epoch-run
7Grok 4.6xAI99.2%Epoch-run
8Claude Opus 5Anthropic98.9%Epoch-run
9DeepSeek V4 ProDeepSeek98.6%Epoch-run
10Claude Opus 4.8Anthropic98.3%Epoch-run
11GPT-5.6 LunaOpenAI98.3%Epoch-run
12GPT-5.4OpenAI97.8%Epoch-run
13Claude Opus 4.7Anthropic97.8%Epoch-run
14Grok 4.5xAI97.8%Epoch-run
15Kimi K3Moonshot AI97.2%Epoch-run
16Gemini 3.7 FlashGoogle97.2%Epoch-run
17GPT-5.2OpenAI96.1%Epoch-run
18Kimi K2.6Moonshot AI96.1%Epoch-run
19Gemini 3.1 ProGoogle95.6%Epoch-run
20Gemini 3.5 FlashGoogle95.6%Epoch-run
21Gemini 3 FlashGoogle95.6%Epoch-run
22Kimi K2.7 CodeMoonshot AI95.6%Epoch-run
23Qwen3.7 MaxAlibaba95.6%Epoch-run
24Claude Sonnet 5Anthropic94.7%Epoch-run
25Claude Opus 4.6Anthropic94.4%Epoch-run
26DeepSeek V4 FlashDeepSeek94.4%Epoch-run
27Gemini 3.6 FlashGoogle94.2%Epoch-run
28GLM 5.3 FlashZhipu AI93.9%Epoch-run
29GLM 5.1Zhipu AI93.3%Epoch-run
30Qwen3.6 PlusAlibaba93.3%Epoch-run
31Qwen3.7 PlusAlibaba93.3%Epoch-run
32Grok 4.3xAI93.3%Epoch-run
33Grok 4.20xAI92.2%Epoch-run
34Kimi K2.5Moonshot AI92.2%Epoch-run
35GPT-5OpenAI91.4%Epoch-run
36Gemini 3 ProGoogle91.4%Epoch-run
37Qwen3.6 MaxAlibaba91.1%Epoch-run
38GLM 5.3Zhipu AI91.1%Epoch-run
39Qwen3.6 27BAlibaba91.1%Epoch-run
40Inkling SmallThinking Machines90.0%Epoch-run
41GPT-5.4 miniOpenAI88.9%Epoch-run
42InklingThinking Machines88.9%Epoch-run
43Muse SparkMeta88.9%Epoch-run
44Qwen3.5 397B-A17BAlibaba88.9%Epoch-run
45GPT-OSS 120BOpenAI88.9%Epoch-run
46GPT-5.1OpenAI88.6%Epoch-run
47DeepSeek V3.2DeepSeek87.8%Epoch-run
48GPT-5.4 nanoOpenAI87.8%Epoch-run
49Qwen3.7 FlashAlibaba86.7%Epoch-run
50Qwen3.5 PlusAlibaba86.7%Epoch-run
51Nemotron 3 UltraNVIDIA86.7%Epoch-run
52GPT-5 miniOpenAI86.7%Epoch-run
53Qwen3.6 35B-A3BAlibaba86.7%Epoch-run
54GLM 5.2Zhipu AI86.4%Epoch-run
55Claude Opus 4.5Anthropic86.1%Epoch-run
56Claude Sonnet 4.6Anthropic85.8%Epoch-run
57GPT-5.5OpenAI84.4%Epoch-run
58Qwen3.6 FlashAlibaba84.4%Epoch-run
59Qwen3.5 FlashAlibaba84.4%Epoch-run
60Gemini 2.5 ProGoogle84.2%Epoch-run
61Grok 4xAI84.0%Epoch-run
62GLM 4.7Zhipu AI83.3%Epoch-run
63Kimi K2Moonshot AI83.1%Epoch-run
64Gemma 4 26B A4BGoogle82.2%Epoch-run
65GPT-5 nanoOpenAI81.1%Epoch-run
66GLM 5Zhipu AI80.0%Epoch-run
67Gemini 3.1 Flash-LiteGoogle80.0%Epoch-run
68Claude Sonnet 4.5Anthropic77.8%Epoch-run
69Qwen3 MaxAlibaba73.3%Epoch-run
70Gemma 4 31BGoogle73.3%Epoch-run
71MiniMax M3MiniMax71.1%Epoch-run
72Gemini 3.5 Flash-LiteGoogle71.1%Epoch-run
73Qwen3.5 35B-A3BAlibaba70.0%Epoch-run
74Claude Opus 4.1Anthropic68.9%Epoch-run
75GPT-5.5 InstantOpenAI68.1%Epoch-run
76Claude Haiku 4.5Anthropic66.7%Epoch-run
77GPT-OSS 20BOpenAI65.3%Epoch-run
78Qwen3.5 9BAlibaba61.7%Epoch-run
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.