All benchmarks
CodingHard set

SWE-bench Verified

500 human-validated GitHub issues the model must resolve in the real repository so the hidden tests pass. Epoch AI runs every model in one standardized scaffold.

Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.

Read with care. One shared scaffold for every model; labs' own numbers with bespoke scaffolds can run 10 points higher and are not used here.

Publisher: Epoch AI (internal runs)

What a model is asked to do

Given a real GitHub issue and its repository, write a patch that makes the project's hidden tests pass.

For exampleIssue: the CSV exporter drops the header row when the first column is empty. Find the cause in the codebase and fix it without breaking existing exports.

Why it matters. This is the daily work of a software engineer: understanding an unfamiliar codebase and fixing the right thing.

The example is original and illustrative, not an item from the dataset.

Models scored here
26
Who produced the numbers
Run by Epoch AI
Items graded
n = 500 (0.2% each)
Best published result
83.5%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (83.5%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Epoch AI ran the evaluation itself, under settings it documents per model, and publishes the log of every answer.

Snapshot September 8, 2026

Ranking on this test

26 models on SWE-bench Verified

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.2 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1Claude Opus 4.7Anthropic83.5%Epoch-run
2Gemini 3.5 FlashGoogle79.3%Epoch-run
3Claude Opus 4.6Anthropic78.7%Epoch-run
4GLM 5.2Zhipu AI78.7%Epoch-run
5DeepSeek V4 ProDeepSeek77.6%Epoch-run
6Qwen3.7 MaxAlibaba77.3%Epoch-run
7GPT-5.4OpenAI76.9%Epoch-run
8Claude Opus 4.5Anthropic76.7%Epoch-run
9Qwen3.6 MaxAlibaba76.7%Epoch-run
10Kimi K2.6Moonshot AI76.7%Epoch-run
11Gemini 3.1 ProGoogle75.6%Epoch-run
12Gemini 3 FlashGoogle75.4%Epoch-run
13Claude Sonnet 4.6Anthropic75.2%Epoch-run
14GPT-5.3 CodexOpenAI74.8%Epoch-run
15GLM 5.1Zhipu AI74.2%Epoch-run
16GPT-5.2OpenAI73.8%Epoch-run
17Kimi K2.5Moonshot AI73.8%Epoch-run
18GPT-5OpenAI73.6%Epoch-run
19Claude Opus 4.1Anthropic73.3%Epoch-run
20Gemini 3 ProGoogle72.9%Epoch-run
21GLM 5Zhipu AI72.1%Epoch-run
22Claude Sonnet 4.5Anthropic71.3%Epoch-run
23GPT-5.1OpenAI68.0%Epoch-run
24GPT-5 miniOpenAI64.7%Epoch-run
25Qwen3.6 PlusAlibaba57.9%Epoch-run
26Gemini 2.5 ProGoogle57.6%Epoch-run
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.