SciCode
Research-grade scientific programming: 338 sub-problems from 80 research problems across physics, chemistry, biology, materials and math that must pass unit tests. As run by Artificial Analysis.
Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.
Publisher: SciCode / Artificial AnalysisWhat a model is asked to do
Write research-grade scientific code, one sub-problem at a time, that passes unit tests.
For exampleImplement a numerical solver for a one-dimensional Schrodinger equation with the given potential, then use it to compute the first three energy levels.
Why it matters. Scientific programming needs the math and the code to be right at the same time.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 73
- Who produced the numbers
- Official leaderboard
- Items graded
- n = 338 (0.3% each)
- Best published result
- 62.0%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (62.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
73 models on SciCode
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.3 points, so read gaps smaller than that as noise.

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.