All benchmarks
Coding

SciCode

Research-grade scientific programming: 338 sub-problems from 80 research problems across physics, chemistry, biology, materials and math that must pass unit tests. As run by Artificial Analysis.

Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.

Publisher: SciCode / Artificial Analysis

What a model is asked to do

Write research-grade scientific code, one sub-problem at a time, that passes unit tests.

For exampleImplement a numerical solver for a one-dimensional Schrodinger equation with the given potential, then use it to compute the first three energy levels.

Why it matters. Scientific programming needs the math and the code to be right at the same time.

The example is original and illustrative, not an item from the dataset.

Models scored here
73
Who produced the numbers
Official leaderboard
Items graded
n = 338 (0.3% each)
Best published result
62.0%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (62.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

73 models on SciCode

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.3 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1Claude Fable 5.1Anthropic62.0%Leaderboard
2Claude Fable 5Anthropic60.2%Leaderboard
3Gemini 3.1 ProGoogle58.9%Leaderboard
4Kimi K3Moonshot AI58.7%Leaderboard
5Muse Spark 1.1Meta58.2%Leaderboard
6Gemini 3.7 FlashGoogle57.9%Leaderboard
7GPT-5.6 SolOpenAI56.9%Leaderboard
8GPT-5.4OpenAI56.6%Leaderboard
9GPT-6 AstraOpenAI56.5%Leaderboard
10GLM 5.3Zhipu AI56.5%Leaderboard
11Muse Spark 1.2Meta56.4%Leaderboard
12GPT-5.5OpenAI56.1%Leaderboard
13Claude Opus 5Anthropic55.7%Leaderboard
14Grok 4.6xAI54.6%Leaderboard
15Claude Opus 4.7Anthropic54.5%Leaderboard
16Gemini 3.8 FlashGoogle54.4%Leaderboard
17Grok 4.5xAI54.1%Leaderboard
18GPT-5.6 TerraOpenAI53.9%Leaderboard
19Claude Sonnet 5Anthropic53.6%Leaderboard
20Claude Opus 4.8Anthropic53.5%Leaderboard
21Kimi K2.6Moonshot AI53.5%Leaderboard
22Gemini 3.5 FlashGoogle53.1%Leaderboard
23Qwen3.8 MaxAlibaba52.9%Leaderboard
24Gemini 3.6 FlashGoogle52.7%Leaderboard
25GPT-5.6 LunaOpenAI52.5%Leaderboard
26Muse SparkMeta51.5%Leaderboard
27GLM 5.2Zhipu AI50.5%Leaderboard
28MiMo V2.5 ProXiaomi50.2%Leaderboard
29DeepSeek V4 ProDeepSeek50.0%Leaderboard
30DeepSeek V4 FlashDeepSeek49.9%Leaderboard
31GPT-5.4 miniOpenAI49.9%Leaderboard
32Kimi K2.5Moonshot AI49.0%Leaderboard
33Qwen3.7 MaxAlibaba48.8%Leaderboard
34Inkling SmallThinking Machines48.7%Leaderboard
35GPT-5.5 InstantOpenAI48.6%Leaderboard
36Kimi K2.7 CodeMoonshot AI47.5%Leaderboard
37Grok 4.3xAI47.3%Leaderboard
38MiniMax M2.7MiniMax47.0%Leaderboard
39GPT-5.4 nanoOpenAI46.9%Leaderboard
40Claude Sonnet 4.6Anthropic46.8%Leaderboard
41InklingThinking Machines46.1%Leaderboard
42GLM 5.3 FlashZhipu AI46.1%Leaderboard
43Qwen3.7 PlusAlibaba45.5%Leaderboard
44MiniMax M3MiniMax45.4%Leaderboard
45GLM 4.7Zhipu AI45.1%Leaderboard
46Claude Sonnet 4.5Anthropic44.7%Leaderboard
47Qwen3.8 27BAlibaba44.7%Leaderboard
48GLM 5.1Zhipu AI43.8%Leaderboard
49Gemma 4 31BGoogle43.4%Leaderboard
50GPT-5.1OpenAI43.3%Leaderboard
51Claude Haiku 4.5Anthropic43.3%Leaderboard
52MiMo V2.5Xiaomi43.1%Leaderboard
53GPT-5OpenAI42.9%Leaderboard
54Gemini 2.5 ProGoogle42.8%Leaderboard
55Gemini 3.1 Flash-LiteGoogle41.9%Leaderboard
56Gemini 3.5 Flash-LiteGoogle40.9%Leaderboard
57Qwen3.6 PlusAlibaba40.7%Leaderboard
58DeepSeek V3.1 TerminusDeepSeek40.6%Leaderboard
59Gemma 4 26B A4BGoogle40.0%Leaderboard
60Nemotron 3 UltraNVIDIA39.9%Leaderboard
61Mistral Medium 3.5Mistral AI39.6%Leaderboard
62GPT-5 miniOpenAI39.2%Leaderboard
63DeepSeek V3.2DeepSeek38.9%Leaderboard
64GPT-OSS 120BOpenAI38.9%Leaderboard
65GLM 4.6Zhipu AI38.4%Leaderboard
66Qwen3.6 27BAlibaba37.3%Leaderboard
67Mistral Large 3Mistral AI36.2%Leaderboard
68Nemotron 3 SuperNVIDIA36.0%Leaderboard
69Qwen3.6 35B-A3BAlibaba35.8%Leaderboard
70Qwen3.5 122B-A10BAlibaba35.6%Leaderboard
71GPT-OSS 20BOpenAI34.4%Leaderboard
72Qwen3.5 35B-A3BAlibaba29.3%Leaderboard
73Qwen3.5 9BAlibaba27.5%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.