Terminal-Bench
Real tasks done in a terminal — set up servers, fix builds, wrangle data — checked by tests. Best published agent harness per model.
Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.
Read with care. Harnesses differ by model, so this compares model-plus-harness systems, not models alone.
Publisher: Terminal-BenchWhat a model is asked to do
Complete a real task inside a terminal, such as setting up a service, fixing a build or wrangling data, verified by tests.
For exampleThe build fails after a dependency upgrade. Diagnose it from the logs, pin what needs pinning, and get the test suite green.
Why it matters. The terminal is where agents do real operations work. This shows whether one can be left alone with a shell.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 40
- Who produced the numbers
- Official leaderboard
- Best published result
- 84.7%
- License
- Apache 2.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (84.7%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
40 models on Terminal-Bench
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | GPT-5.5OpenAI | 84.7% | 100.0 | agent: NexAU-AHE | Leaderboard |
| 2 | GPT-5.4OpenAI | 81.8% | 96.6 | 2026-03-05 · agent: ForgeCode | Leaderboard |
| 3 | Claude Opus 4.7Anthropic | 80.2% | 94.7 | agent: WOZCODE | Leaderboard |
| 4 | Gemini 3.1 ProGoogle | 80.2% | 94.7 | preview · agent: TongAgents | Leaderboard |
| 5 | Claude Opus 4.6Anthropic | 79.8% | 94.2 | agent: ForgeCode | Leaderboard |
| 6 | GPT-5.3 CodexOpenAI | 78.4% | 92.6 | agent: SageAgent | Leaderboard |
| 7 | Gemini 3 ProGoogle | 69.4% | 81.9 | preview · agent: Ante | Leaderboard |
| 8 | GPT-5.2 CodexOpenAI | 66.5% | 78.5 | agent: Deep Agents | Leaderboard |
| 9 | GPT-5.2OpenAI | 64.9% | 76.6 | 2025-12-11 · medium · agent: Droid | Leaderboard |
| 10 | Gemini 3 FlashGoogle | 64.3% | 75.9 | preview · agent: Junie CLI | Leaderboard |
| 11 | Claude Opus 4.5Anthropic | 63.1% | 74.5 | 20251101 · agent: Droid | Leaderboard |
| 12 | GPT-5.1 Codex MiniOpenAI | 61.6% | 72.7 | agent: hookele | Leaderboard |
| 13 | GPT-5.1 CodexOpenAI | 60.4% | 71.3 | max · agent: Codex CLI | Leaderboard |
| 14 | Grok 4.20xAI | 57.3% | 67.7 | agent: grok-cli | Leaderboard |
| 15 | Claude Sonnet 4.6Anthropic | 53.4% | 63.0 | agent: Simplai Agent | Leaderboard |
| 16 | GLM 5Zhipu AI | 52.4% | 61.9 | agent: Terminus 2 | Leaderboard |
| 17 | GPT-5OpenAI | 49.6% | 58.6 | 2025-08-07 · agent: Codex CLI | Leaderboard |
| 18 | GPT-5.1OpenAI | 47.6% | 56.2 | 2025-11-13 · agent: Terminus 2 | Leaderboard |
| 19 | Claude Sonnet 4.5Anthropic | 46.5% | 54.9 | 20250929 · agent: CAMEL-AI | Leaderboard |
| 20 | MiniMax M2.7MiniMax | 45.1% | 53.2 | agent: IndusAGI Coding Agent | Leaderboard |
| 21 | GPT-5 CodexOpenAI | 44.3% | 52.3 | agent: Codex CLI | Leaderboard |
| 22 | Kimi K2.5Moonshot AI | 43.2% | 51.0 | agent: Terminus 2 | Leaderboard |
| 23 | MiniMax M2.5MiniMax | 42.7% | 50.4 | agent: cchuter | Leaderboard |
| 24 | DeepSeek V3.2DeepSeek | 39.6% | 46.8 | agent: Terminus 2 | Leaderboard |
| 25 | Claude Opus 4.1Anthropic | 38.0% | 44.9 | 20250805 · agent: Terminus 2 | Leaderboard |
| 26 | MiniMax M2.1MiniMax | 36.6% | 43.2 | agent: Crux | Leaderboard |
| 27 | Kimi K2Moonshot AI | 35.7% | 42.1 | thinking · agent: Terminus 2 | Leaderboard |
| 28 | Claude Haiku 4.5Anthropic | 35.5% | 41.9 | 20251001 · agent: Goose | Leaderboard |
| 29 | GPT-5 miniOpenAI | 34.8% | 41.1 | 2025-08-07 · agent: spoox-m | Leaderboard |
| 30 | GLM 4.7Zhipu AI | 33.4% | 39.4 | agent: Terminus 2 | Leaderboard |
| 31 | Gemini 2.5 ProGoogle | 32.6% | 38.5 | agent: Terminus 2 | Leaderboard |
| 32 | MiniMax M2MiniMax | 30.0% | 35.4 | agent: Terminus 2 | Leaderboard |
| 33 | Grok 4xAI | 27.2% | 32.1 | agent: OpenHands | Leaderboard |
| 34 | GLM 4.6Zhipu AI | 24.5% | 28.9 | agent: Terminus 2 | Leaderboard |
| 35 | Qwen3.6 35B-A3BAlibaba | 23.0% | 27.2 | agent: little-coder | Leaderboard |
| 36 | GPT-5 nanoOpenAI | 21.8% | 25.7 | 2025-08-07 · agent: spoox-o-m | Leaderboard |
| 37 | GPT-OSS 120BOpenAI | 18.7% | 22.1 | agent: Terminus 2 | Leaderboard |
| 38 | Gemini 2.5 FlashGoogle | 17.1% | 20.2 | agent: Mini-SWE-Agent | Leaderboard |
| 39 | Qwen3.5 9BAlibaba | 9.2% | 10.9 | agent: little-coder | Leaderboard |
| 40 | GPT-OSS 20BOpenAI | 3.4% | 4.0 | agent: Mini-SWE-Agent | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.