All benchmarks
Human preferenceBeside the Index

Text Arena

Millions of blind head-to-head votes from real users. Style-controlled Arena score — the market's human-preference standard.

Blind head-to-head votes from real users on real prompts. Human preference is a single-test category in this snapshot (fewer than 2 comparable tests), so it is shown beside the Index and never averaged into it.

Publisher: Arena

What a model is asked to do

Answer the same prompt as a rival model; real users vote blind on which answer they prefer.

For exampleRewrite this dense lease clause for a first-time renter without losing the two conditions that actually matter.

Why it matters. Millions of blind votes on real prompts are the market's measure of what people prefer.

The example is original and illustrative, not an item from the dataset.

Models scored here
92
Who produced the numbers
Official leaderboard
Best published result
1,507
License
CC BY 4.0
Direction
Higher is better
Scale
Published as an Arena rating from blind head-to-head votes. For the Index, each rating becomes the expected win rate against the board leader under the Elo model, scaled so parity reads 100: 100 points behind the leader is a 36% win rate and reads 72. The scale does not depend on which weak model happens to be listed.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

92 models on Text Arena

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, as an expected win rate against the board leader). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1Claude Fable 5Anthropic1,507Leaderboard
2Claude Opus 4.6Anthropic1,505Leaderboard
3Claude Fable 5.1Anthropic1,504Leaderboard
4Claude Opus 4.7Anthropic1,502Leaderboard
5Muse Spark 1.2Meta1,499Leaderboard
6Gemini 3.8 FlashGoogle1,494Leaderboard
7Claude Opus 5Anthropic1,493Leaderboard
8Muse Spark 1.1Meta1,492Leaderboard
9Gemini 3.7 FlashGoogle1,491Leaderboard
10Kimi K3Moonshot AI1,489Leaderboard
11Muse SparkMeta1,488Leaderboard
12Gemini 3.1 ProGoogle1,487Leaderboard
13Gemini 3 ProGoogle1,486Leaderboard
14GPT-5.6 SolOpenAI1,483Leaderboard
15Claude Opus 4.8Anthropic1,482Leaderboard
16GPT-5.5OpenAI1,482Leaderboard
17GLM 5.3Zhipu AI1,482Leaderboard
18Gemini 3.6 FlashGoogle1,480Leaderboard
19Qwen3.8 MaxAlibaba1,480Leaderboard
20Gemini 3.5 FlashGoogle1,479Leaderboard
21GPT-5.4OpenAI1,477Leaderboard
22Grok 4.20xAI1,475Leaderboard
23GLM 5.3 FlashZhipu AI1,474Leaderboard
24Gemini 3 FlashGoogle1,474Leaderboard
25Qwen3.7 MaxAlibaba1,474Leaderboard
26GPT-5.5 InstantOpenAI1,474Leaderboard
27Claude Opus 4.5Anthropic1,473Leaderboard
28Claude Sonnet 4.6Anthropic1,472Leaderboard
29GLM 5.2Zhipu AI1,472Leaderboard
30Grok 4.5xAI1,471Leaderboard
31ERNIE 5.1Baidu1,468Leaderboard
32MiMo V2.5 ProXiaomi1,468Leaderboard
33GPT-5.6 TerraOpenAI1,466Leaderboard
34GLM 5.1Zhipu AI1,466Leaderboard
35Grok 4.1xAI1,465Leaderboard
36Claude Sonnet 5Anthropic1,462Leaderboard
37Kimi K2.6Moonshot AI1,461Leaderboard
38Qwen3.6 MaxAlibaba1,460Leaderboard
39DeepSeek V4 ProDeepSeek1,460Leaderboard
40GLM 5Zhipu AI1,458Leaderboard
41Gemini 3.5 Flash-LiteGoogle1,457Leaderboard
42Claude Sonnet 4.5Anthropic1,456Leaderboard
43Seed 2.0 ProByteDance1,456Leaderboard
44Hunyuan 3Tencent1,455Leaderboard
45Qwen3.7 PlusAlibaba1,455Leaderboard
46GPT-5.1OpenAI1,455Leaderboard
47GPT-5.6 LunaOpenAI1,453Leaderboard
48Gemma 4 31BGoogle1,451Leaderboard
49Kimi K2.5Moonshot AI1,451Leaderboard
50Claude Opus 4.1Anthropic1,450Leaderboard
51GPT-5.4 miniOpenAI1,448Leaderboard
52Gemini 2.5 ProGoogle1,446Leaderboard
53Qwen3.6 PlusAlibaba1,444Leaderboard
54MiniMax M3MiniMax1,443Leaderboard
55Grok 4.3xAI1,443Leaderboard
56GLM 4.7Zhipu AI1,442Leaderboard
57Qwen3.5 397B-A17BAlibaba1,441Leaderboard
58InklingThinking Machines1,439Leaderboard
59DeepSeek V4 FlashDeepSeek1,438Leaderboard
60Gemma 4 26B A4BGoogle1,438Leaderboard
61GPT-5.2OpenAI1,438Leaderboard
62Qwen3.8 27BAlibaba1,436Leaderboard
63GPT-5OpenAI1,434Leaderboard
64Qwen3 MaxAlibaba1,435Leaderboard
65MiMo V2.5Xiaomi1,434Leaderboard
66Gemini 3.1 Flash-LiteGoogle1,432Leaderboard
67Kimi K2Moonshot AI1,430Leaderboard
68Grok 4.1 FastxAI1,430Leaderboard
69Mistral Medium 3.5Mistral AI1,427Leaderboard
70Nemotron 3 UltraNVIDIA1,426Leaderboard
71DeepSeek V3.2DeepSeek1,425Leaderboard
72GLM 4.6Zhipu AI1,425Leaderboard
73Grok 4 FastxAI1,418Leaderboard
74DeepSeek V3.1 TerminusDeepSeek1,418Leaderboard
75Qwen3.5 122B-A10BAlibaba1,417Leaderboard
76MiniMax M2.7MiniMax1,415Leaderboard
77Mistral Large 3Mistral AI1,414Leaderboard
78Claude Haiku 4.5Anthropic1,413Leaderboard
79Grok 4xAI1,411Leaderboard
80Gemini 2.5 FlashGoogle1,410Leaderboard
81Qwen3.5 27BAlibaba1,408Leaderboard
82Inkling SmallThinking Machines1,407Leaderboard
83GPT-5.4 nanoOpenAI1,402Leaderboard
84Qwen3.5 FlashAlibaba1,397Leaderboard
85Qwen3.5 35B-A3BAlibaba1,395Leaderboard
86MiniMax M2.5MiniMax1,391Leaderboard
87GPT-5 miniOpenAI1,389Leaderboard
88MiniMax M2.1MiniMax1,384Leaderboard
89GPT-OSS 120BOpenAI1,352Leaderboard
90MiniMax M2MiniMax1,346Leaderboard
91GPT-5 nanoOpenAI1,337Leaderboard
92GPT-OSS 20BOpenAI1,317Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.