Vision Arena
Blind head-to-head votes on prompts that include an image. Style-controlled Arena score.
Understanding images alongside text, judged by human preference. Multimodal is a single-test category in this snapshot (fewer than 2 comparable tests), so it is shown beside the Index and never averaged into it.
Publisher: ArenaWhat a model is asked to do
Answer prompts that include an image; real users vote blind between two models' answers.
For exampleHere is a photo of a circuit board after a power surge. Which component most likely failed, and what points to it?
Why it matters. Seeing is half of most real tasks: screenshots, documents, photos and charts.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 54
- Who produced the numbers
- Official leaderboard
- Best published result
- 1,313
- License
- CC BY 4.0
- Direction
- Higher is better
- Scale
- Published as an Arena rating from blind head-to-head votes. For the Index, each rating becomes the expected win rate against the board leader under the Elo model, scaled so parity reads 100: 100 points behind the leader is a 36% win rate and reads 72. The scale does not depend on which weak model happens to be listed.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
54 models on Vision Arena
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, as an expected win rate against the board leader). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | Claude Fable 5Anthropic | 1,313 | 100.0 | board 2026-08-27 | Leaderboard |
| 2 | Claude Opus 4.7Anthropic | 1,301 | 96.7 | high · board 2026-08-27 | Leaderboard |
| 3 | Qwen3.8 MaxAlibaba | 1,300 | 96.2 | max · board 2026-08-27 | Leaderboard |
| 4 | Claude Opus 4.6Anthropic | 1,299 | 95.9 | high · board 2026-08-27 | Leaderboard |
| 5 | Muse SparkMeta | 1,294 | 94.4 | board 2026-08-27 | Leaderboard |
| 6 | Muse Spark 1.2Meta | 1,292 | 93.9 | xhigh · board 2026-08-27 | Leaderboard |
| 7 | Claude Opus 5Anthropic | 1,290 | 93.5 | high · board 2026-08-27 | Leaderboard |
| 8 | Gemini 3 ProGoogle | 1,289 | 93.0 | board 2026-08-27 | Leaderboard |
| 9 | GPT-5.5OpenAI | 1,286 | 92.3 | board 2026-08-27 | Leaderboard |
| 10 | Claude Opus 4.8Anthropic | 1,285 | 92.0 | high · board 2026-08-27 | Leaderboard |
| 11 | Gemini 3.6 FlashGoogle | 1,285 | 91.9 | high · board 2026-08-27 | Leaderboard |
| 12 | Gemini 3.5 FlashGoogle | 1,284 | 91.7 | high · board 2026-08-27 | Leaderboard |
| 13 | GPT-5.6 SolOpenAI | 1,282 | 91.2 | xhigh · board 2026-08-27 | Leaderboard |
| 14 | Grok 4.5xAI | 1,282 | 91.1 | board 2026-08-27 | Leaderboard |
| 15 | GPT-5.4OpenAI | 1,281 | 90.7 | high · board 2026-08-27 | Leaderboard |
| 16 | Muse Spark 1.1Meta | 1,279 | 90.4 | board 2026-08-27 | Leaderboard |
| 17 | Gemini 3.1 ProGoogle | 1,278 | 90.1 | preview · board 2026-08-27 | Leaderboard |
| 18 | GPT-5.5 InstantOpenAI | 1,278 | 89.8 | board 2026-08-27 | Leaderboard |
| 19 | Claude Sonnet 4.6Anthropic | 1,275 | 89.2 | board 2026-08-27 | Leaderboard |
| 20 | Gemini 3 FlashGoogle | 1,272 | 88.4 | board 2026-08-27 | Leaderboard |
| 21 | GLM 5.3 FlashZhipu AI | 1,273 | 88.4 | board 2026-08-27 | Leaderboard |
| 22 | Claude Sonnet 5Anthropic | 1,267 | 87.0 | high · board 2026-08-27 | Leaderboard |
| 23 | Gemini 3.5 Flash-LiteGoogle | 1,266 | 86.7 | board 2026-08-27 | Leaderboard |
| 24 | GPT-5.6 TerraOpenAI | 1,266 | 86.5 | xhigh · board 2026-08-27 | Leaderboard |
| 25 | Qwen3.7 PlusAlibaba | 1,266 | 86.5 | board 2026-08-27 | Leaderboard |
| 26 | Kimi K2.6Moonshot AI | 1,263 | 85.6 | board 2026-08-27 | Leaderboard |
| 27 | Gemma 4 31BGoogle | 1,261 | 85.2 | board 2026-08-27 | Leaderboard |
| 28 | Seed 2.0 ProByteDance | 1,257 | 84.0 | board 2026-08-27 | Leaderboard |
| 29 | Grok 4.20xAI | 1,256 | 83.7 | reasoning · board 2026-08-27 | Leaderboard |
| 30 | GPT-5.6 LunaOpenAI | 1,254 | 83.0 | xhigh · board 2026-08-27 | Leaderboard |
| 31 | GPT-5.4 miniOpenAI | 1,252 | 82.6 | high · board 2026-08-27 | Leaderboard |
| 32 | Qwen3.8 27BAlibaba | 1,251 | 82.5 | board 2026-08-27 | Leaderboard |
| 33 | GPT-5.1OpenAI | 1,250 | 82.0 | high · board 2026-08-27 | Leaderboard |
| 34 | Kimi K2.5Moonshot AI | 1,250 | 81.9 | thinking · board 2026-08-27 | Leaderboard |
| 35 | Qwen3.5 397B-A17BAlibaba | 1,247 | 81.3 | board 2026-08-27 | Leaderboard |
| 36 | Gemini 2.5 ProGoogle | 1,246 | 81.1 | board 2026-08-27 | Leaderboard |
| 37 | GPT-5.2OpenAI | 1,244 | 80.4 | high · board 2026-08-27 | Leaderboard |
| 38 | Gemma 4 26B A4BGoogle | 1,242 | 79.7 | board 2026-08-27 | Leaderboard |
| 39 | Grok 4.3xAI | 1,241 | 79.6 | board 2026-08-27 | Leaderboard |
| 40 | MiniMax M3MiniMax | 1,237 | 78.5 | board 2026-08-27 | Leaderboard |
| 41 | MiMo V2.5Xiaomi | 1,235 | 77.9 | board 2026-08-27 | Leaderboard |
| 42 | Gemini 3.1 Flash-LiteGoogle | 1,235 | 77.9 | preview · board 2026-08-27 | Leaderboard |
| 43 | Qwen3.5 122B-A10BAlibaba | 1,227 | 75.6 | board 2026-08-27 | Leaderboard |
| 44 | Qwen3.5 27BAlibaba | 1,219 | 73.5 | board 2026-08-27 | Leaderboard |
| 45 | Gemini 2.5 FlashGoogle | 1,214 | 72.4 | board 2026-08-27 | Leaderboard |
| 46 | GPT-5OpenAI | 1,210 | 71.3 | high · board 2026-08-27 | Leaderboard |
| 47 | Inkling SmallThinking Machines | 1,207 | 70.3 | board 2026-08-27 | Leaderboard |
| 48 | Mistral Large 3Mistral AI | 1,205 | 69.9 | board 2026-08-27 | Leaderboard |
| 49 | GPT-5.4 nanoOpenAI | 1,201 | 68.9 | high · board 2026-08-27 | Leaderboard |
| 50 | Mistral Medium 3.5Mistral AI | 1,199 | 68.3 | board 2026-08-27 | Leaderboard |
| 51 | Grok 4.1 FastxAI | 1,195 | 67.2 | reasoning · board 2026-08-27 | Leaderboard |
| 52 | Grok 4xAI | 1,184 | 64.4 | board 2026-08-27 | Leaderboard |
| 53 | GPT-5 miniOpenAI | 1,182 | 64.0 | high · board 2026-08-27 | Leaderboard |
| 54 | GPT-5 nanoOpenAI | 1,145 | 55.1 | high · board 2026-08-27 | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.