All benchmarks
AgenticHard set

GDPval

Economically valuable work across 44 occupations, judged by industry professionals against deliverables from their peers. Win rate on the 220-task gold set.

Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.

Read with care. Authored and graded by OpenAI, which also competes on it. Treat as a lab-run leaderboard.

Publisher: OpenAI (external evaluations)

What a model is asked to do

Produce a real work deliverable across 44 occupations, judged against what industry professionals hand in.

For examplePrepare a quarterly cash-flow forecast for a regional dental practice from the attached statements, listing every assumption.

Why it matters. Judged against professionals, so it measures economic value rather than trivia.

The example is original and illustrative, not an item from the dataset.

Models scored here
8
Who produced the numbers
Official leaderboard
Items graded
n = 220 (0.5% each)
Best published result
49.7%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (49.7%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

8 models on GDPval

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.5 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1GPT-5.2OpenAI49.7%Leaderboard
2Claude Opus 4.5Anthropic45.5%Leaderboard
3Claude Opus 4.1Anthropic43.6%Leaderboard
4Claude Sonnet 4.5Anthropic42.5%Leaderboard
5Gemini 3 ProGoogle40.3%Leaderboard
6GPT-5OpenAI34.8%Leaderboard
7Gemini 2.5 ProGoogle23.3%Leaderboard
8Grok 4xAI21.1%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.