Code Arena (WebDev)
Blind head-to-head votes on web apps two models built from the same prompt. Arena score.
Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.
Publisher: ArenaWhat a model is asked to do
Build a web app from the same prompt as a rival model; real users vote blind on the result.
For exampleBuild a kanban board with drag-and-drop, a dark mode and local persistence, all in a single page.
Why it matters. People judging finished apps side by side is the most honest measure of front-end quality.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 86
- Who produced the numbers
- Official leaderboard
- Best published result
- 1,797
- License
- CC BY 4.0
- Direction
- Higher is better
- Scale
- Published as an Arena rating from blind head-to-head votes. For the Index, each rating becomes the expected win rate against the board leader under the Elo model, scaled so parity reads 100: 100 points behind the leader is a 36% win rate and reads 72. The scale does not depend on which weak model happens to be listed.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
86 models on Code Arena (WebDev)
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, as an expected win rate against the board leader). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | GPT-6 AstraOpenAI | 1,797 | 100.0 | max · board 2026-09-05 | Leaderboard |
| 2 | Claude Fable 5.1Anthropic | 1,762 | 90.0 | max · board 2026-09-05 | Leaderboard |
| 3 | Claude Opus 5Anthropic | 1,688 | 69.5 | max · board 2026-09-05 | Leaderboard |
| 4 | Qwen3.8 MaxAlibaba | 1,686 | 69.1 | board 2026-09-05 | Leaderboard |
| 5 | Kimi K3Moonshot AI | 1,674 | 66.1 | max · board 2026-09-05 | Leaderboard |
| 6 | Claude Fable 5Anthropic | 1,629 | 55.1 | board 2026-09-05 | Leaderboard |
| 7 | Qwen3.8 Flash NextAlibaba | 1,626 | 54.5 | board 2026-09-05 | Leaderboard |
| 8 | Grok 4.6xAI | 1,625 | 54.1 | high · board 2026-09-05 | Leaderboard |
| 9 | Muse Spark 1.3Meta | 1,622 | 53.6 | xhigh · board 2026-09-05 | Leaderboard |
| 10 | GPT-5.6 SolOpenAI | 1,617 | 52.5 | xhigh · codex-harness · board 2026-09-05 | Leaderboard |
| 11 | GLM 5.3Zhipu AI | 1,609 | 50.6 | max · board 2026-09-05 | Leaderboard |
| 12 | GLM 5.3 FlashZhipu AI | 1,605 | 49.7 | board 2026-09-05 | Leaderboard |
| 13 | Qwen3.8 27BAlibaba | 1,594 | 47.4 | board 2026-09-05 | Leaderboard |
| 14 | Gemini 3.7 FlashGoogle | 1,587 | 46.0 | high · board 2026-09-05 | Leaderboard |
| 15 | GLM 5.2Zhipu AI | 1,587 | 45.9 | max · board 2026-09-05 | Leaderboard |
| 16 | DeepSeek V4 ProDeepSeek | 1,582 | 44.9 | high-20260813 · board 2026-09-05 | Leaderboard |
| 17 | DeepSeek V4 FlashDeepSeek | 1,580 | 44.5 | high · board 2026-09-05 | Leaderboard |
| 18 | Gemini 3.8 FlashGoogle | 1,567 | 42.1 | high · board 2026-09-05 | Leaderboard |
| 19 | Claude Opus 4.8Anthropic | 1,563 | 41.2 | high · board 2026-09-05 | Leaderboard |
| 20 | Claude Opus 4.7Anthropic | 1,557 | 40.2 | board 2026-09-05 | Leaderboard |
| 21 | Grok 4.5xAI | 1,556 | 39.9 | board 2026-09-05 | Leaderboard |
| 22 | Claude Opus 4.6Anthropic | 1,546 | 38.2 | high · board 2026-09-05 | Leaderboard |
| 23 | Muse Spark 1.1Meta | 1,541 | 37.3 | board 2026-09-05 | Leaderboard |
| 24 | Gemini 3.6 FlashGoogle | 1,538 | 36.8 | high · board 2026-09-05 | Leaderboard |
| 25 | Claude Sonnet 5Anthropic | 1,537 | 36.6 | high · board 2026-09-05 | Leaderboard |
| 26 | Muse Spark 1.2Meta | 1,534 | 36.1 | xhigh · board 2026-09-05 | Leaderboard |
| 27 | Claude Sonnet 4.6Anthropic | 1,521 | 34.0 | board 2026-09-05 | Leaderboard |
| 28 | GPT-5.6 TerraOpenAI | 1,520 | 33.7 | xhigh · codex-harness · board 2026-09-05 | Leaderboard |
| 29 | GPT-5.6 LunaOpenAI | 1,519 | 33.7 | xhigh · codex-harness · board 2026-09-05 | Leaderboard |
| 30 | Qwen3.7 MaxAlibaba | 1,517 | 33.3 | max-20260517 · board 2026-09-05 | Leaderboard |
| 31 | Hunyuan 3Tencent | 1,512 | 32.4 | board 2026-09-05 | Leaderboard |
| 32 | GPT-5.5OpenAI | 1,510 | 32.1 | xhigh · codex-harness · board 2026-09-05 | Leaderboard |
| 33 | Kimi K2.6Moonshot AI | 1,509 | 32.0 | board 2026-09-05 | Leaderboard |
| 34 | GLM 5.1Zhipu AI | 1,508 | 31.9 | board 2026-09-05 | Leaderboard |
| 35 | Gemini 3.5 FlashGoogle | 1,500 | 30.7 | high · board 2026-09-05 | Leaderboard |
| 36 | Claude Opus 4.5Anthropic | 1,495 | 29.8 | 20251101-high-32k · board 2026-09-05 | Leaderboard |
| 37 | MiniMax M3MiniMax | 1,487 | 28.8 | board 2026-09-05 | Leaderboard |
| 38 | Qwen3.6 MaxAlibaba | 1,479 | 27.6 | max-preview · board 2026-09-05 | Leaderboard |
| 39 | MiMo V2.5 ProXiaomi | 1,475 | 27.1 | board 2026-09-05 | Leaderboard |
| 40 | Kimi K2.7 CodeMoonshot AI | 1,472 | 26.7 | board 2026-09-05 | Leaderboard |
| 41 | GPT-5.4OpenAI | 1,463 | 25.5 | high · codex-harness · board 2026-09-05 | Leaderboard |
| 42 | Qwen3.6 PlusAlibaba | 1,460 | 25.1 | board 2026-09-05 | Leaderboard |
| 43 | Gemini 3.5 Flash-LiteGoogle | 1,449 | 23.8 | board 2026-09-05 | Leaderboard |
| 44 | Gemini 3.1 ProGoogle | 1,446 | 23.4 | preview · board 2026-09-05 | Leaderboard |
| 45 | Gemini 3 ProGoogle | 1,439 | 22.5 | board 2026-09-05 | Leaderboard |
| 46 | Gemini 3 FlashGoogle | 1,438 | 22.4 | board 2026-09-05 | Leaderboard |
| 47 | MiMo V2.5Xiaomi | 1,438 | 22.4 | board 2026-09-05 | Leaderboard |
| 48 | Kimi K2.5Moonshot AI | 1,436 | 22.3 | thinking · board 2026-09-05 | Leaderboard |
| 49 | GLM 5Zhipu AI | 1,436 | 22.2 | board 2026-09-05 | Leaderboard |
| 50 | GLM 4.7Zhipu AI | 1,434 | 22.1 | board 2026-09-05 | Leaderboard |
| 51 | GPT-5OpenAI | 1,419 | 20.4 | medium · board 2026-09-05 | Leaderboard |
| 52 | GPT-5.2OpenAI | 1,416 | 20.1 | board 2026-09-05 | Leaderboard |
| 53 | InklingThinking Machines | 1,409 | 19.3 | board 2026-09-05 | Leaderboard |
| 54 | GPT-5.3 CodexOpenAI | 1,409 | 19.3 | codex-harness · board 2026-09-05 | Leaderboard |
| 55 | Inkling SmallThinking Machines | 1,405 | 19.0 | board 2026-09-05 | Leaderboard |
| 56 | Qwen3.5 397B-A17BAlibaba | 1,399 | 18.3 | board 2026-09-05 | Leaderboard |
| 57 | MiniMax M2.7MiniMax | 1,398 | 18.3 | board 2026-09-05 | Leaderboard |
| 58 | GPT-5.4 miniOpenAI | 1,397 | 18.2 | high · board 2026-09-05 | Leaderboard |
| 59 | Claude Sonnet 4.5Anthropic | 1,392 | 17.7 | 20250929-high-32k · board 2026-09-05 | Leaderboard |
| 60 | GPT-5.1OpenAI | 1,392 | 17.7 | medium · board 2026-09-05 | Leaderboard |
| 61 | Claude Opus 4.1Anthropic | 1,389 | 17.4 | 20250805 · board 2026-09-05 | Leaderboard |
| 62 | MiniMax M2.1MiniMax | 1,387 | 17.3 | preview · board 2026-09-05 | Leaderboard |
| 63 | MiniMax M2.5MiniMax | 1,384 | 17.0 | board 2026-09-05 | Leaderboard |
| 64 | Grok 4.20xAI | 1,374 | 16.1 | reasoning · board 2026-09-05 | Leaderboard |
| 65 | Gemma 4 31BGoogle | 1,363 | 15.2 | board 2026-09-05 | Leaderboard |
| 66 | Gemma 4 26B A4BGoogle | 1,361 | 15.0 | board 2026-09-05 | Leaderboard |
| 67 | DeepSeek V3.2DeepSeek | 1,360 | 15.0 | thinking · board 2026-09-05 | Leaderboard |
| 68 | Qwen3.5 122B-A10BAlibaba | 1,358 | 14.8 | board 2026-09-05 | Leaderboard |
| 69 | Qwen3.5 27BAlibaba | 1,357 | 14.7 | board 2026-09-05 | Leaderboard |
| 70 | Grok 4.3xAI | 1,356 | 14.7 | board 2026-09-05 | Leaderboard |
| 71 | GLM 4.6Zhipu AI | 1,340 | 13.5 | board 2026-09-05 | Leaderboard |
| 72 | GPT-5.2 CodexOpenAI | 1,338 | 13.3 | board 2026-09-05 | Leaderboard |
| 73 | GPT-5.1 CodexOpenAI | 1,336 | 13.1 | board 2026-09-05 | Leaderboard |
| 74 | Claude Haiku 4.5Anthropic | 1,329 | 12.7 | 20251001 · board 2026-09-05 | Leaderboard |
| 75 | Kimi K2Moonshot AI | 1,322 | 12.2 | board 2026-09-05 | Leaderboard |
| 76 | MiniMax M2MiniMax | 1,298 | 10.7 | board 2026-09-05 | Leaderboard |
| 77 | Mistral Medium 3.5Mistral AI | 1,265 | 8.9 | board 2026-09-05 | Leaderboard |
| 78 | Gemini 3.1 Flash-LiteGoogle | 1,254 | 8.4 | preview · board 2026-09-05 | Leaderboard |
| 79 | Qwen3.5 35B-A3BAlibaba | 1,250 | 8.2 | board 2026-09-05 | Leaderboard |
| 80 | GPT-5.1 Codex MiniOpenAI | 1,244 | 7.9 | board 2026-09-05 | Leaderboard |
| 81 | Grok 4.1 FastxAI | 1,240 | 7.8 | reasoning · board 2026-09-05 | Leaderboard |
| 82 | Qwen3.5 FlashAlibaba | 1,238 | 7.7 | board 2026-09-05 | Leaderboard |
| 83 | Mistral Large 3Mistral AI | 1,229 | 7.3 | board 2026-09-05 | Leaderboard |
| 84 | Gemini 2.5 ProGoogle | 1,226 | 7.2 | board 2026-09-05 | Leaderboard |
| 85 | Grok 4.1xAI | 1,211 | 6.6 | thinking · board 2026-09-05 | Leaderboard |
| 86 | Grok 4 FastxAI | 1,162 | 5.0 | reasoning · board 2026-09-05 | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.