All benchmarks
ReasoningInternally runnable

ARC-AGI-1

The original Abstraction and Reasoning Corpus — grid puzzles solved from a handful of examples. Semi-private evaluation set, verified by ARC Prize.

Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.

Read with care. Close to saturated at the frontier; scores depend on the compute budget a lab chose.

Publisher: ARC Prize Foundation

What a model is asked to do

Solve the original Abstraction and Reasoning Corpus: small grids, a handful of examples, one hidden rule.

For exampleTwo examples show a pattern being extended along its row until it reaches the border. Complete the third grid.

Why it matters. The first ARC set is close to saturated at the frontier, so it now shows whether smaller models can generalize at all.

The example is original and illustrative, not an item from the dataset.

Models scored here
51
Who produced the numbers
Official leaderboard
Items graded
n = 100 (1.0% each)
Best published result
98.5%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (98.5%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

51 models on ARC-AGI-1

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 1.0 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1GPT-6 AstraOpenAI98.5%Leaderboard
2Claude Fable 5Anthropic98.5%Leaderboard
3Gemini 3.1 ProGoogle98.0%Leaderboard
4Claude Fable 5.1Anthropic97.5%Leaderboard
5Claude Opus 5Anthropic97.5%Leaderboard
6GPT-5.6 SolOpenAI97.5%Leaderboard
7GPT-5.6 TerraOpenAI96.5%Leaderboard
8GPT-5.5 ProOpenAI96.5%Leaderboard
9Gemini 3 Deep ThinkGoogle96.0%Leaderboard
10Gemini 3.7 FlashGoogle95.5%Leaderboard
11GPT-5.5OpenAI95.0%Leaderboard
12Kimi K3Moonshot AI94.5%Leaderboard
13GPT-5.4 ProOpenAI94.5%Leaderboard
14Claude Opus 4.6Anthropic94.0%Leaderboard
15GPT-5.4OpenAI93.7%Leaderboard
16Claude Opus 4.7Anthropic93.5%Leaderboard
17Claude Opus 4.8Anthropic92.5%Leaderboard
18Gemini 3.5 FlashGoogle92.5%Leaderboard
19Gemini 3.6 FlashGoogle91.2%Leaderboard
20DeepSeek V4 ProDeepSeek90.5%Leaderboard
21GPT-5.2 ProOpenAI90.5%Leaderboard
22Grok 4.20xAI89.5%Leaderboard
23DeepSeek V4 FlashDeepSeek89.0%Leaderboard
24GPT-5.6 LunaOpenAI88.0%Leaderboard
25Grok 4.6xAI87.5%Leaderboard
26Grok 4.5xAI87.2%Leaderboard
27Claude Sonnet 4.6Anthropic86.5%Leaderboard
28GPT-5.2OpenAI86.2%Leaderboard
29Gemini 3 FlashGoogle84.7%Leaderboard
30Inkling SmallThinking Machines84.0%Leaderboard
31Claude Opus 4.5Anthropic80.0%Leaderboard
32InklingThinking Machines79.5%Leaderboard
33GLM 5.2Zhipu AI77.0%Leaderboard
34Gemini 3 ProGoogle75.0%Leaderboard
35GPT-5.1OpenAI72.8%Leaderboard
36GPT-5 ProOpenAI70.2%Leaderboard
37Grok 4xAI66.7%Leaderboard
38GPT-5OpenAI65.7%Leaderboard
39Kimi K2.5Moonshot AI65.3%Leaderboard
40Claude Sonnet 4.5Anthropic63.7%Leaderboard
41GPT-5.4 miniOpenAI63.7%Leaderboard
42MiniMax M2.5MiniMax63.7%Leaderboard
43DeepSeek V3.2DeepSeek57.0%Leaderboard
44GPT-5 miniOpenAI54.3%Leaderboard
45Gemini 3.5 Flash-LiteGoogle53.5%Leaderboard
46GPT-5.4 nanoOpenAI51.5%Leaderboard
47Grok 4 FastxAI48.5%Leaderboard
48Claude Haiku 4.5Anthropic47.7%Leaderboard
49GLM 5Zhipu AI44.7%Leaderboard
50Gemini 2.5 ProGoogle41.0%Leaderboard
51GPT-5 nanoOpenAI20.7%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.