All benchmarks
AgenticHard set

Terminal-Bench

Real tasks done in a terminal — set up servers, fix builds, wrangle data — checked by tests. Best published agent harness per model.

Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.

Read with care. Harnesses differ by model, so this compares model-plus-harness systems, not models alone.

Publisher: Terminal-Bench

What a model is asked to do

Complete a real task inside a terminal, such as setting up a service, fixing a build or wrangling data, verified by tests.

For exampleThe build fails after a dependency upgrade. Diagnose it from the logs, pin what needs pinning, and get the test suite green.

Why it matters. The terminal is where agents do real operations work. This shows whether one can be left alone with a shell.

The example is original and illustrative, not an item from the dataset.

Models scored here
40
Who produced the numbers
Official leaderboard
Best published result
84.7%
License
Apache 2.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (84.7%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

40 models on Terminal-Bench

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1GPT-5.5OpenAI84.7%Leaderboard
2GPT-5.4OpenAI81.8%Leaderboard
3Claude Opus 4.7Anthropic80.2%Leaderboard
4Gemini 3.1 ProGoogle80.2%Leaderboard
5Claude Opus 4.6Anthropic79.8%Leaderboard
6GPT-5.3 CodexOpenAI78.4%Leaderboard
7Gemini 3 ProGoogle69.4%Leaderboard
8GPT-5.2 CodexOpenAI66.5%Leaderboard
9GPT-5.2OpenAI64.9%Leaderboard
10Gemini 3 FlashGoogle64.3%Leaderboard
11Claude Opus 4.5Anthropic63.1%Leaderboard
12GPT-5.1 Codex MiniOpenAI61.6%Leaderboard
13GPT-5.1 CodexOpenAI60.4%Leaderboard
14Grok 4.20xAI57.3%Leaderboard
15Claude Sonnet 4.6Anthropic53.4%Leaderboard
16GLM 5Zhipu AI52.4%Leaderboard
17GPT-5OpenAI49.6%Leaderboard
18GPT-5.1OpenAI47.6%Leaderboard
19Claude Sonnet 4.5Anthropic46.5%Leaderboard
20MiniMax M2.7MiniMax45.1%Leaderboard
21GPT-5 CodexOpenAI44.3%Leaderboard
22Kimi K2.5Moonshot AI43.2%Leaderboard
23MiniMax M2.5MiniMax42.7%Leaderboard
24DeepSeek V3.2DeepSeek39.6%Leaderboard
25Claude Opus 4.1Anthropic38.0%Leaderboard
26MiniMax M2.1MiniMax36.6%Leaderboard
27Kimi K2Moonshot AI35.7%Leaderboard
28Claude Haiku 4.5Anthropic35.5%Leaderboard
29GPT-5 miniOpenAI34.8%Leaderboard
30GLM 4.7Zhipu AI33.4%Leaderboard
31Gemini 2.5 ProGoogle32.6%Leaderboard
32MiniMax M2MiniMax30.0%Leaderboard
33Grok 4xAI27.2%Leaderboard
34GLM 4.6Zhipu AI24.5%Leaderboard
35Qwen3.6 35B-A3BAlibaba23.0%Leaderboard
36GPT-5 nanoOpenAI21.8%Leaderboard
37GPT-OSS 120BOpenAI18.7%Leaderboard
38Gemini 2.5 FlashGoogle17.1%Leaderboard
39Qwen3.5 9BAlibaba9.2%Leaderboard
40GPT-OSS 20BOpenAI3.4%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.