Complete software tasks of increasing length; the score is the task length a model finishes with 50% reliability.
3.7 h
77.8% of the best (17.4 h) on a log scale
#8 of 14
max
METR publishes wide confidence intervals around each horizon; the point estimate is used here.
Source: METR