ARC-AGI-2
Abstract visual reasoning puzzles that are easy for people and designed to resist memorization — the benchmark built to measure general fluid intelligence. Semi-private evaluation set, verified by ARC Prize.
Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.
Read with care. Scores depend on the compute budget a lab chose; ARC Prize publishes cost per task beside every score and this ranking does not.
Publisher: ARC Prize FoundationWhat a model is asked to do
Infer the hidden rule from a few input and output grid pairs, then apply it to a new grid.
For exampleThree examples show colored shapes being reflected across a diagonal and recolored by their size. Produce the output for a fourth, unseen input.
Why it matters. Novel visual puzzles resist memorization, so they measure fluid intelligence rather than recall.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 51
- Who produced the numbers
- Official leaderboard
- Items graded
- n = 120 (0.8% each)
- Best published result
- 95.0%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (95.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
51 models on ARC-AGI-2
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.8 points, so read gaps smaller than that as noise.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 95.0% | 100.0 | max | Leaderboard |
| 2 | GPT-5.6 SolOpenAI | 92.5% | 97.4 | max | Leaderboard |
| 3 | Claude Opus 5Anthropic | 90.4% | 95.2 | max | Leaderboard |
| 4 | Claude Fable 5.1Anthropic | 90.0% | 94.7 | max | Leaderboard |
| 5 | Claude Fable 5Anthropic | 89.2% | 93.9 | max | Leaderboard |
| 6 | GPT-5.5OpenAI | 85.0% | 89.5 | xhigh | Leaderboard |
| 7 | Gemini 3.7 FlashGoogle | 84.6% | 89.0 | high | Leaderboard |
| 8 | GPT-5.5 ProOpenAI | 84.6% | 89.0 | high | Leaderboard |
| 9 | Gemini 3 Deep ThinkGoogle | 84.6% | 89.0 | preview | Leaderboard |
| 10 | GPT-5.6 TerraOpenAI | 83.9% | 88.3 | max | Leaderboard |
| 11 | GPT-5.4 ProOpenAI | 83.3% | 87.7 | 2026-03-05 · xhigh | Leaderboard |
| 12 | Gemini 3.1 ProGoogle | 77.1% | 81.2 | preview | Leaderboard |
| 13 | Claude Opus 4.7Anthropic | 75.8% | 79.8 | max | Leaderboard |
| 14 | GPT-5.4OpenAI | 74.0% | 77.8 | 2026-03-05 · xhigh | Leaderboard |
| 15 | Claude Opus 4.8Anthropic | 72.1% | 75.9 | high | Leaderboard |
| 16 | Gemini 3.5 FlashGoogle | 72.1% | 75.9 | high | Leaderboard |
| 17 | Claude Opus 4.6Anthropic | 69.2% | 72.8 | 120k | Leaderboard |
| 18 | Grok 4.6xAI | 67.1% | 70.6 | xhigh | Leaderboard |
| 19 | Grok 4.20xAI | 65.1% | 68.6 | — | Leaderboard |
| 20 | DeepSeek V4 FlashDeepSeek | 61.4% | 64.6 | max | Leaderboard |
| 21 | DeepSeek V4 ProDeepSeek | 61.3% | 64.5 | max | Leaderboard |
| 22 | Kimi K3Moonshot AI | 60.4% | 63.6 | max | Leaderboard |
| 23 | Claude Sonnet 4.6Anthropic | 60.4% | 63.6 | high | Leaderboard |
| 24 | Gemini 3.6 FlashGoogle | 60.4% | 63.6 | high | Leaderboard |
| 25 | GPT-5.6 LunaOpenAI | 59.5% | 62.7 | max | Leaderboard |
| 26 | GPT-5.2 ProOpenAI | 54.2% | 57.0 | 2025-12-11 · high | Leaderboard |
| 27 | GPT-5.2OpenAI | 52.9% | 55.7 | 2025-12-11 · xhigh | Leaderboard |
| 28 | Grok 4.5xAI | 52.6% | 55.4 | high | Leaderboard |
| 29 | Inkling SmallThinking Machines | 40.1% | 42.3 | xhigh | Leaderboard |
| 30 | Claude Opus 4.5Anthropic | 37.6% | 39.6 | 20251101 · 64k | Leaderboard |
| 31 | InklingThinking Machines | 36.5% | 38.5 | — | Leaderboard |
| 32 | Gemini 3 FlashGoogle | 33.6% | 35.4 | preview | Leaderboard |
| 33 | Gemini 3 ProGoogle | 31.1% | 32.7 | preview | Leaderboard |
| 34 | GLM 5.2Zhipu AI | 22.8% | 24.0 | — | Leaderboard |
| 35 | GPT-5.4 miniOpenAI | 18.9% | 19.9 | 2026-03-17 · xhigh | Leaderboard |
| 36 | GPT-5 ProOpenAI | 18.3% | 19.3 | 2025-10-06 | Leaderboard |
| 37 | GPT-5.1OpenAI | 17.6% | 18.6 | 2025-11-13 · high | Leaderboard |
| 38 | Grok 4xAI | 16.0% | 16.8 | — | Leaderboard |
| 39 | Claude Sonnet 4.5Anthropic | 13.6% | 14.3 | 20250929 · 32k | Leaderboard |
| 40 | Kimi K2.5Moonshot AI | 11.8% | 12.4 | — | Leaderboard |
| 41 | Gemini 3.5 Flash-LiteGoogle | 10.3% | 10.8 | high | Leaderboard |
| 42 | GPT-5OpenAI | 9.9% | 10.4 | 2025-08-07 · high | Leaderboard |
| 43 | GPT-5.4 nanoOpenAI | 5.7% | 6.0 | 2026-03-17 · xhigh | Leaderboard |
| 44 | Grok 4 FastxAI | 5.3% | 5.6 | — | Leaderboard |
| 45 | Gemini 2.5 ProGoogle | 4.9% | 5.1 | 32k | Leaderboard |
| 46 | GLM 5Zhipu AI | 4.9% | 5.1 | — | Leaderboard |
| 47 | MiniMax M2.5MiniMax | 4.9% | 5.1 | — | Leaderboard |
| 48 | GPT-5 miniOpenAI | 4.4% | 4.7 | 2025-08-07 · high | Leaderboard |
| 49 | DeepSeek V3.2DeepSeek | 4.0% | 4.2 | — | Leaderboard |
| 50 | Claude Haiku 4.5Anthropic | 4.0% | 4.2 | 20251001 · 32k | Leaderboard |
| 51 | GPT-5 nanoOpenAI | 2.6% | 2.7 | 2025-08-07 · high | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.