Mock AIME 2024–2025
Competition mathematics from the OTIS Mock AIME sets — 45 problems, exact-answer graded, run 16 times by Epoch AI.
Competition and research mathematics, graded on the final answer or the proof. One of 5 comparable tests in math, a category the Index averages.
Read with care. 45 problems: one problem is 2.2 points, so gaps of a few points are within noise.
Publisher: Epoch AI (internal runs)What a model is asked to do
Solve competition mathematics problems that have a single exact integer answer.
For exampleFind the number of ordered pairs of positive integers (a, b) with a + b = 2026 such that a and b share no prime factor.
Why it matters. Olympiad-style problems need long, careful chains of reasoning, and there is no partial credit.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 78
- Who produced the numbers
- Run by Epoch AI
- Items graded
- n = 45 (2.2% each)
- Best published result
- 100.0%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (100.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.
Snapshot September 8, 2026
Ranking on this test
78 models on Mock AIME 2024–2025
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 2.2 points, so read gaps smaller than that as noise.

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.