CursorBench
Real coding-agent requests drawn from everyday editor sessions, scored against the change the developer actually shipped.
Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.
Read with care. Published by Cursor from its own product traffic and harness; not independently reproduced.
Publisher: CursorWhat a model is asked to do
Handle a real coding-agent request from an everyday editor session, judged against the change the developer actually shipped.
For exampleMake this module's configuration loading asynchronous and update every call site, keeping the behavior identical.
Why it matters. Judged against what a person shipped, so it measures usefulness inside a real workflow.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 21
- Who produced the numbers
- Official leaderboard
- Best published result
- 73.4%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (73.4%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
21 models on CursorBench
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | Claude Fable 5.1Anthropic | 73.4% | 100.0 | max · reasoning level: Max | Leaderboard |
| 2 | Grok 4.6xAI | 70.8% | 96.5 | xhigh · reasoning level: Extra High | Leaderboard |
| 3 | Claude Fable 5Anthropic | 70.5% | 96.0 | max · reasoning level: Max | Leaderboard |
| 4 | Claude Opus 5Anthropic | 70.0% | 95.4 | max · reasoning level: Max | Leaderboard |
| 5 | Gemini 3.8 FlashGoogle | 69.2% | 94.3 | high · reasoning level: High | Leaderboard |
| 6 | GPT-5.6 SolOpenAI | 67.2% | 91.6 | max · reasoning level: Max | Leaderboard |
| 7 | GPT-5.6 TerraOpenAI | 64.9% | 88.4 | max · reasoning level: Max | Leaderboard |
| 8 | Claude Opus 4.7Anthropic | 64.8% | 88.3 | max · reasoning level: Max | Leaderboard |
| 9 | Claude Opus 4.8Anthropic | 62.3% | 84.9 | max · reasoning level: Max | Leaderboard |
| 10 | Gemini 3.7 FlashGoogle | 61.6% | 83.9 | high · reasoning level: High | Leaderboard |
| 11 | Claude Sonnet 5Anthropic | 61.5% | 83.8 | max · reasoning level: Max | Leaderboard |
| 12 | GPT-5.6 LunaOpenAI | 61.1% | 83.2 | max · reasoning level: Max | Leaderboard |
| 13 | Kimi K3Moonshot AI | 60.8% | 82.8 | max · reasoning level: Max | Leaderboard |
| 14 | GPT-5.5OpenAI | 58.4% | 79.6 | xhigh · reasoning level: Extra High | Leaderboard |
| 15 | GLM 5.2Zhipu AI | 55.0% | 74.9 | max · reasoning level: Max | Leaderboard |
| 16 | Gemini 3.6 FlashGoogle | 53.5% | 72.9 | high · reasoning level: High | Leaderboard |
| 17 | Gemini 3.5 FlashGoogle | 49.8% | 67.8 | — | Leaderboard |
| 18 | Kimi K2.7 CodeMoonshot AI | 49.7% | 67.7 | — | Leaderboard |
| 19 | Claude Sonnet 4.6Anthropic | 49.0% | 66.8 | max · reasoning level: Max | Leaderboard |
| 20 | Kimi K2.6Moonshot AI | 47.6% | 64.9 | — | Leaderboard |
| 21 | Kimi K2.5Moonshot AI | 31.9% | 43.5 | — | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.