All benchmarks
Coding

FrontierCode

Hard, realistic software tasks run through a coding-agent harness. Main score.

Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.

Read with care. Published by Cognition, a coding-agent vendor, from its own harness; not independently reproduced.

Publisher: Cognition

What a model is asked to do

Complete a hard, realistic software task through a coding-agent harness.

For exampleAdd end-to-end encryption to the attachment path of this messaging service, including key rotation, with the existing tests still passing.

Why it matters. Multi-file, multi-hour engineering tasks are where agentic coding is heading.

The example is original and illustrative, not an item from the dataset.

Models scored here
27
Who produced the numbers
Official leaderboard
Best published result
53.5%
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (53.5%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

27 models on FrontierCode

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1Claude Fable 5Anthropic53.5%Leaderboard
2Claude Opus 5Anthropic53.4%Leaderboard
3GPT-6 AstraOpenAI53.3%Leaderboard
4Claude Fable 5.1Anthropic50.9%Leaderboard
5Grok 4.6xAI48.0%Leaderboard
6GPT-5.6 SolOpenAI47.5%Leaderboard
7Claude Opus 4.8Anthropic46.5%Leaderboard
8Kimi K3Moonshot AI44.2%Leaderboard
9Gemini 3.7 FlashGoogle43.6%Leaderboard
10GPT-5.5OpenAI43.0%Leaderboard
11Claude Sonnet 5Anthropic42.7%Leaderboard
12Grok 4.5xAI42.4%Leaderboard
13GPT-5.6 TerraOpenAI41.3%Leaderboard
14GPT-5.6 LunaOpenAI39.8%Leaderboard
15Claude Opus 4.7Anthropic38.5%Leaderboard
16Gemini 3.6 FlashGoogle34.4%Leaderboard
17Kimi K2.7 CodeMoonshot AI30.1%Leaderboard
18GPT-5.4 miniOpenAI27.0%Leaderboard
19Claude Opus 4.6Anthropic26.6%Leaderboard
20GLM 5.2Zhipu AI24.5%Leaderboard
21Claude Sonnet 4.6Anthropic24.3%Leaderboard
22DeepSeek V4 FlashDeepSeek18.8%Leaderboard
23DeepSeek V4 ProDeepSeek17.6%Leaderboard
24MiniMax M3MiniMax14.7%Leaderboard
25InklingThinking Machines14.0%Leaderboard
26Qwen3.7 PlusAlibaba10.2%Leaderboard
27Mistral Medium 3.5Mistral AI8.0%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.