All benchmarks
Agentic

OSWorld 2.0

108 long-horizon computer-use workflows across real desktop and web applications; a skilled person takes a median of about 1.6 hours per task. Share of tasks completed in full at the 500-step budget.

Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.

Read with care. A separate, far harder benchmark from OSWorld-Verified: binary completion is low for every model, and the leaderboard also reports a partial-credit score.

Publisher: XLANG Lab

What a model is asked to do

Complete a multi-hour workflow across real desktop and web applications, end to end, with the environment changing as you work.

For exampleReconcile this month's expense claims: pull the receipts from the mailbox, match them against the bank export in the spreadsheet, flag the two that do not add up, and submit the rest through the reimbursement portal.

Why it matters. Computer use is how an agent reaches every tool that has no API, and long workflows are where agents lose the thread.

The example is original and illustrative, not an item from the dataset.

Models scored here
9
Who produced the numbers
Official leaderboard
Items graded
n = 108 (0.9% each)
Best published result
31.4%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (31.4%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

9 models on OSWorld 2.0

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.9 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1Claude Opus 5Anthropic31.4%Leaderboard
2GPT-5.6 SolOpenAI27.3%Leaderboard
3Claude Opus 4.8Anthropic20.6%Leaderboard
4Claude Opus 4.7Anthropic18.2%Leaderboard
5GPT-5.5OpenAI13.0%Leaderboard
6Claude Sonnet 4.6Anthropic9.3%Leaderboard
7Kimi K2.6Moonshot AI4.6%Leaderboard
8MiniMax M3MiniMax4.6%Leaderboard
9Qwen3.7 PlusAlibaba2.8%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.