METR Time Horizon
The length of software task (in human-expert minutes) a model completes with 50% reliability. Shown in minutes; normalized on a log scale from one minute to the frontier.
Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.
Read with care. METR publishes wide confidence intervals around each horizon; the point estimate is used here.
Publisher: METRWhat a model is asked to do
Complete software tasks of increasing length; the score is the task length a model finishes with 50% reliability.
For exampleTasks run from a five-minute regex fix to a multi-hour feature build with its own test plan.
Why it matters. The horizon a model can work through unsupervised is the clearest measure of how much you can delegate.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 14
- Who produced the numbers
- Official leaderboard
- Best published result
- 17.4 h
- Log-scale floor
- 1 min
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- An open-ended value. For the Index, each score is placed on a log scale from a fixed floor (1 min) to the best published result (17.4 h), so the frontier reads 100 and doubling counts the same anywhere on the scale.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
14 models on METR Time Horizon
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, on a log scale). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | Claude Opus 4.6Anthropic | 12.0 h | 94.6 | — | Leaderboard |
| 2 | Gemini 3.1 ProGoogle | 6.4 h | 85.6 | preview | Leaderboard |
| 3 | GPT-5.2OpenAI | 5.9 h | 84.4 | 2025-12-11 · high | Leaderboard |
| 4 | GPT-5.3 CodexOpenAI | 5.8 h | 84.2 | — | Leaderboard |
| 5 | GPT-5.4OpenAI | 5.7 h | 83.9 | 2026-03-05 | Leaderboard |
| 6 | Claude Opus 4.5Anthropic | 4.9 h | 81.7 | 20251101 | Leaderboard |
| 7 | Gemini 3 ProGoogle | 3.7 h | 77.9 | preview | Leaderboard |
| 8 | GPT-5.1 CodexOpenAI | 3.7 h | 77.8 | max | Leaderboard |
| 9 | GPT-5OpenAI | 3.4 h | 76.4 | 2025-08-07 · high | Leaderboard |
| 10 | Claude Sonnet 4.5Anthropic | 2.0 h | 69.1 | 20250929 · 16k | Leaderboard |
| 11 | Claude Opus 4.1Anthropic | 114 min | 68.1 | 20250805 · 16k | Leaderboard |
| 12 | Grok 4xAI | 110 min | 67.6 | — | Leaderboard |
| 13 | Kimi K2Moonshot AI | 54 min | 57.4 | thinking | Leaderboard |
| 14 | GPT-OSS 120BOpenAI | 42 min | 53.8 | — | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.