All benchmarks
Reasoning

LiveBench

A contamination-resistant benchmark refreshed monthly — reasoning, coding, math, data analysis, language and instruction following. Global average.

Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.

Read with care. Epoch's copy of this leaderboard lags the live site, so few current models have a score here.

Publisher: LiveBench

What a model is asked to do

Answer a monthly-refreshed mix of reasoning, coding, math, data analysis, language and instruction-following questions with objective answers.

For exampleGiven this table of quarterly sales, return exactly the JSON the question asks for, with the one column that needs a unit conversion corrected.

Why it matters. Fresh questions every month mean the answers cannot have leaked into training data.

The example is original and illustrative, not an item from the dataset.

Models scored here
1
Who produced the numbers
Official leaderboard
Best published result
82.3%
License
Apache 2.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (82.3%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

1 model on LiveBench

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. With only 1 model scored, this test is shown but not yet counted toward the Index.

#ModelScoreProduced by
1GPT-5.1OpenAI78.8%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.