CritPt
Unpublished research-level physics problems graded by an official server — accuracy on the full set, as run by Artificial Analysis.
Hard, novel problems: graduate science, abstract puzzles, expert exams. One of 6 comparable tests in reasoning, a category the Index averages.
Publisher: CritPt / Artificial AnalysisWhat a model is asked to do
Solve an unpublished research-level physics problem to a numeric or symbolic answer checked by an official server.
For exampleDerive the leading correction to the decay rate of a metastable state coupled to a bath with the given spectral density, and give it in closed form.
Why it matters. Research physics is far past textbook recall. Only a few models produce anything a physicist would accept.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 78
- Who produced the numbers
- Official leaderboard
- Best published result
- 32.3%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (32.3%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
78 models on CritPt
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.