All benchmarks
Math

ProofBench

Full written proofs, not final answers, graded for rigor.

Competition and research mathematics, graded on the final answer or the proof. One of 5 comparable tests in math, a category the Index averages.

Publisher: Vals AI

What a model is asked to do

Write a complete, rigorous proof, graded for rigor rather than for a final answer.

For exampleProve that every sequence of real numbers has a monotone subsequence, and say exactly where completeness is or is not used.

Why it matters. A proof shows the reasoning itself. A right answer with a wrong argument scores nothing.

The example is original and illustrative, not an item from the dataset.

Models scored here
61
Who produced the numbers
Official leaderboard
Best published result
100.0%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (100.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

61 models on ProofBench

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1Claude Fable 5.1Anthropic100.0%Leaderboard
2GPT-6 AstraOpenAI99.0%Leaderboard
3Claude Opus 5Anthropic99.0%Leaderboard
4Claude Fable 5Anthropic95.0%Leaderboard
5Kimi K3Moonshot AI87.0%Leaderboard
6GPT-5.6 SolOpenAI83.0%Leaderboard
7Claude Sonnet 5Anthropic77.0%Leaderboard
8GPT-5.6 TerraOpenAI74.0%Leaderboard
9Claude Opus 4.8Anthropic69.0%Leaderboard
10GPT-5.6 LunaOpenAI60.0%Leaderboard
11Gemini 3.7 FlashGoogle58.0%Leaderboard
12Qwen3.8 MaxAlibaba58.0%Leaderboard
13GPT-5.4OpenAI56.0%Leaderboard
14DeepSeek V4 FlashDeepSeek56.0%Leaderboard
15Claude Opus 4.7Anthropic54.0%Leaderboard
16Grok 4.6xAI51.0%Leaderboard
17GPT-5.5OpenAI50.0%Leaderboard
18Claude Opus 4.6Anthropic50.0%Leaderboard
19DeepSeek V4 ProDeepSeek50.0%Leaderboard
20GLM 5.3Zhipu AI49.0%Leaderboard
21Gemini 3.8 FlashGoogle48.0%Leaderboard
22Claude Sonnet 4.6Anthropic45.0%Leaderboard
23Muse Spark 1.2Meta43.0%Leaderboard
24Muse Spark 1.1Meta39.0%Leaderboard
25Gemini 3.6 FlashGoogle36.0%Leaderboard
26Claude Opus 4.5Anthropic36.0%Leaderboard
27GLM 5.2Zhipu AI35.0%Leaderboard
28Gemini 3.5 FlashGoogle31.0%Leaderboard
29Grok 4.5xAI31.0%Leaderboard
30Gemini 3.1 ProGoogle26.0%Leaderboard
31Qwen3.7 MaxAlibaba26.0%Leaderboard
32GLM 5.1Zhipu AI22.2%Leaderboard
33MiMo V2.5 ProXiaomi22.0%Leaderboard
34GPT-5.4 miniOpenAI21.0%Leaderboard
35GLM 5.3 FlashZhipu AI21.0%Leaderboard
36Gemini 3 ProGoogle20.0%Leaderboard
37Claude Sonnet 4.5Anthropic19.0%Leaderboard
38GPT-5OpenAI18.0%Leaderboard
39MiniMax M3MiniMax18.0%Leaderboard
40Muse SparkMeta17.0%Leaderboard
41Kimi K2.6Moonshot AI16.0%Leaderboard
42Qwen3.8 27BAlibaba16.0%Leaderboard
43MiMo V2.5Xiaomi16.0%Leaderboard
44GPT-5.2OpenAI15.0%Leaderboard
45Gemini 3 FlashGoogle15.0%Leaderboard
46Grok 4.20xAI14.0%Leaderboard
47Gemini 3.5 Flash-LiteGoogle13.0%Leaderboard
48GPT-5 nanoOpenAI12.0%Leaderboard
49Grok 4.3xAI11.0%Leaderboard
50GPT-5 miniOpenAI9.0%Leaderboard
51Mistral Medium 3.5Mistral AI9.0%Leaderboard
52GPT-5.1 CodexOpenAI9.0%Leaderboard
53DeepSeek V3.2DeepSeek8.0%Leaderboard
54Inkling SmallThinking Machines6.0%Leaderboard
55GLM 4.7Zhipu AI6.0%Leaderboard
56GPT-5.4 nanoOpenAI5.0%Leaderboard
57Grok 4.1 FastxAI4.0%Leaderboard
58MiniMax M2.5MiniMax4.0%Leaderboard
59MiniMax M2.7MiniMax3.0%Leaderboard
60Nemotron 3 UltraNVIDIA2.0%Leaderboard
61InklingThinking Machines0.0%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.