Complete a multi-hour workflow across real desktop and web applications, end to end, with the environment changing as you work.
31.4%
best published result
of 9 models
max
A separate, far harder benchmark from OSWorld-Verified: binary completion is low for every model, and the leaderboard also reports a partial-credit score.
Source: XLANG Lab