# OpenCharts Benchmarks — AI model rankings from third-party published results > The OpenCharts Index ranks 117 AI models on 31 third-party benchmarks. Every score is expressed relative to the best result its publisher has published on that test (100 = the frontier), averaged inside each capability category, then across the categories that have more than one test. Every number links to the publisher that produced it and carries its provenance. Snapshot: 2026-09-08T17:08:58.223Z. Live page: https://opencharts.com/benchmarks · JSON: https://opencharts.com/benchmarks/rankings.json · Methodology: https://opencharts.com/benchmarks/methodology ## How to read this - Percent scores are the share of the best published result on that test above random guessing (a score at 46% when the best is 46% reads 100), so sitting a harder exam never lowers a model. Arena ratings are the expected win rate against the board leader, scaled so parity reads 100 (100 Elo behind reads 72). Open-ended values (task minutes, dollars) are on a log scale from a fixed published floor to the frontier. The frontier is the publisher's best row, whether or not that model is listed here. A test with fewer than 3 listed models is shown but not counted. - OpenCharts Index = mean of a model's counting category means. Only categories with at least 2 comparable tests are averaged (Reasoning, Coding, Math, Agentic); single-test categories (Knowledge, Multimodal, Human preference, Long context) are reported beside it. A category counts for a model once it has results on at least half of that category's comparable tests. Rank requires ≥ 6 comparable scores across ≥ 3 counting categories; otherwise the model is listed as provisional. - Every Index carries a leave-one-out band: where it lands if any single counted test is dropped. Ranks are printed with the band of places that covers; adjacent models whose bands overlap are statistically tied. - Grade bands (distance from the frontier): A+ ≥ 95, A ≥ 90, A- ≥ 85, B+ ≥ 80, B ≥ 75, B- ≥ 70, C+ ≥ 65, C ≥ 60, C- ≥ 50, D ≥ 40, F ≥ 0. Within 5% of the best published results, on average. - Hard-set score = mean share of the frontier on the hardest, most general tests in the set (Humanity's Last Exam, ARC-AGI-2, SWE-bench Verified, FrontierMath (Tiers 1–3), FrontierMath Tier 4, Terminal-Bench, GDPval, METR Time Horizon), third-party scores only, shown once ≥ 4 exist; models with fewer are listed as pending with the mean they have so far. It measures distance to today's frontier on those tests, nothing more. - Country of origin is the lab's headquarters; every model inherits it. - Verification: "third-party published" means every counted score comes from a publisher other than the model's maker; each score also says whether Epoch AI ran it, an official leaderboard or third-party evaluator produced it, or the developer reported it. "internal" means OpenCharts ran the evaluation itself (first-party engines whose independent verification is TBD); internal models are ranked but never lead a spotlight. ## Ranking #1. GPT-6 Astra (OpenAI, United States) — Index 99.0 (leave-one-out 99.0 to 99.7), grade A+, 14 comparable benchmarks, verification: third-party published. Reasoning 99.6, Knowledge 100.0 (single-test category, shown beside the Index, not averaged), Coding 97.7, Math 99.8, Agentic 98.5 (not counted: fewer than half the category's tests). https://opencharts.com/benchmarks/gpt-6-astra #2. Claude Fable 5.1 (Anthropic, United States) — Index 97.0 (leave-one-out 96.6 to 97.8), grade A+, 17 comparable benchmarks, verification: third-party published. Reasoning 97.5, Knowledge 93.7 (single-test category, shown beside the Index, not averaged), Coding 97.0, Math 96.6, Agentic 88.4 (not counted: fewer than half the category's tests), Human preference 99.2 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-fable-5-1 #3 (3 to 4). Claude Fable 5 (Anthropic, United States) — Index 92.7 (leave-one-out 91.9 to 95.6), grade A, 19 comparable benchmarks, verification: third-party published. Reasoning 93.7, Knowledge 93.5 (single-test category, shown beside the Index, not averaged), Coding 89.4, Math 95.1, Agentic 86.6 (not counted: fewer than half the category's tests), Multimodal 100.0 (single-test category, shown beside the Index, not averaged), Human preference 100.0 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-fable-5 #4 (3 to 4). Claude Opus 5 (Anthropic, United States) — Index 92.6 (leave-one-out 91.7 to 94.3), grade A, 20 comparable benchmarks, verification: third-party published. Reasoning 96.0, Knowledge 79.2 (single-test category, shown beside the Index, not averaged), Coding 90.6, Math 91.1, Agentic 97.3 (not counted: fewer than half the category's tests), Multimodal 93.5 (single-test category, shown beside the Index, not averaged), Human preference 96.0 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-opus-5 #5. GPT-5.6 Sol (OpenAI, United States) — Index 89.8 (leave-one-out 88.7 to 92.4), grade A-, 20 comparable benchmarks, verification: third-party published. Reasoning 94.5, Knowledge 92.2 (single-test category, shown beside the Index, not averaged), Coding 84.0, Math 90.8, Agentic 88.8 (not counted: fewer than half the category's tests), Multimodal 91.2 (single-test category, shown beside the Index, not averaged), Human preference 93.1 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gpt-5-6-sol #6. GPT-5.6 Terra (OpenAI, United States) — Index 81.9 (leave-one-out 80.2 to 85.3), grade B+, 18 comparable benchmarks, verification: third-party published. Reasoning 87.1, Knowledge 57.1 (single-test category, shown beside the Index, not averaged), Coding 74.1, Math 84.5, Agentic 86.5 (not counted: fewer than half the category's tests), Multimodal 86.5 (single-test category, shown beside the Index, not averaged), Human preference 88.3 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gpt-5-6-terra #7 (6 to 9). GPT-5.5 (OpenAI, United States) — Index 80.2 (leave-one-out 78.8 to 82.8), grade B+, 22 comparable benchmarks, verification: third-party published. Reasoning 89.4, Knowledge 83.3 (single-test category, shown beside the Index, not averaged), Coding 74.8, Math 74.9, Agentic 81.5, Multimodal 92.3 (single-test category, shown beside the Index, not averaged), Human preference 92.8 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gpt-5-5 #8 (6 to 10). Kimi K3 (Moonshot AI, China) — Index 79.6 (leave-one-out 77.2 to 83.5), grade B, 18 comparable benchmarks, verification: third-party published. Reasoning 80.5, Knowledge 66.9 (single-test category, shown beside the Index, not averaged), Coding 83.0, Math 75.3, Agentic 79.0 (not counted: fewer than half the category's tests), Human preference 94.7 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/kimi-k3 #9 (7 to 9). Claude Opus 4.8 (Anthropic, United States) — Index 79.5 (leave-one-out 77.7 to 81.8), grade B, 21 comparable benchmarks, verification: third-party published. Reasoning 81.4, Knowledge 70.1 (single-test category, shown beside the Index, not averaged), Coding 77.7, Math 77.6, Agentic 81.2, Multimodal 92.0 (single-test category, shown beside the Index, not averaged), Human preference 92.9 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-opus-4-8 #10 (10 to 11). GPT-5.4 (OpenAI, United States) — Index 77.3 (leave-one-out 75.2 to 79.1), grade B, 21 comparable benchmarks, verification: third-party published. Reasoning 84.0, Knowledge 59.7 (single-test category, shown beside the Index, not averaged), Coding 73.1, Math 72.0, Agentic 80.1, Multimodal 90.7 (single-test category, shown beside the Index, not averaged), Human preference 91.2 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gpt-5-4 #11 (10 to 16). Claude Opus 4.6 (Anthropic, United States) — Index 75.4 (leave-one-out 72.6 to 78.4), grade B, 22 comparable benchmarks, verification: third-party published. Reasoning 83.5, Knowledge 62.2 (single-test category, shown beside the Index, not averaged), Coding 66.6, Math 60.6, Agentic 91.1, Multimodal 95.9 (single-test category, shown beside the Index, not averaged), Human preference 99.3 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-opus-4-6 #12 (9 to 16). Gemini 3.7 Flash (Google, United States) — Index 75.2 (leave-one-out 71.9 to 79.5), grade B, 14 comparable benchmarks, verification: third-party published. Reasoning 82.2, Knowledge 91.5 (single-test category, shown beside the Index, not averaged), Coding 76.2, Math 67.3, Human preference 95.3 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gemini-3-7-flash #13 (10 to 16). Grok 4.6 (xAI, United States) — Index 75.2 (leave-one-out 71.2 to 78.6), grade B, 17 comparable benchmarks, verification: third-party published. Reasoning 82.1, Knowledge 65.2 (single-test category, shown beside the Index, not averaged), Coding 80.2, Math 63.3, Agentic 90.1 (not counted: fewer than half the category's tests). https://opencharts.com/benchmarks/grok-4-6 #14 (10 to 16). Claude Opus 4.7 (Anthropic, United States) — Index 75.1 (leave-one-out 72.4 to 77.8), grade B, 23 comparable benchmarks, verification: third-party published. Reasoning 76.2, Knowledge 68.4 (single-test category, shown beside the Index, not averaged), Coding 78.5, Math 64.8, Agentic 80.9, Multimodal 96.7 (single-test category, shown beside the Index, not averaged), Human preference 98.5 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-opus-4-7 #15 (11 to 17). Gemini 3.1 Pro (Google, United States) — Index 73.4 (leave-one-out 69.9 to 75.7), grade B-, 22 comparable benchmarks, verification: third-party published. Reasoning 88.4, Knowledge 97.2 (single-test category, shown beside the Index, not averaged), Coding 71.6, Math 53.2, Agentic 80.5, Multimodal 90.1 (single-test category, shown beside the Index, not averaged), Human preference 94.1 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gemini-3-1-pro #16 (11 to 17). GPT-5.6 Luna (OpenAI, United States) — Index 72.9 (leave-one-out 70.6 to 75.8), grade B-, 18 comparable benchmarks, verification: third-party published. Reasoning 73.4, Knowledge 54.2 (single-test category, shown beside the Index, not averaged), Coding 68.3, Math 77.1, Agentic 67.7 (not counted: fewer than half the category's tests), Multimodal 83.0 (single-test category, shown beside the Index, not averaged), Human preference 84.4 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gpt-5-6-luna #17 (12 to 18). Claude Sonnet 5 (Anthropic, United States) — Index 71.0 (leave-one-out 68.0 to 75.2), grade B-, 17 comparable benchmarks, verification: third-party published. Reasoning 72.9, Knowledge 44.6 (single-test category, shown beside the Index, not averaged), Coding 72.1, Math 67.9, Agentic 75.3 (not counted: fewer than half the category's tests), Multimodal 87.0 (single-test category, shown beside the Index, not averaged), Human preference 87.2 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-sonnet-5 #18 (17 to 22). Gemini 3.5 Flash (Google, United States) — Index 68.2 (leave-one-out 63.7 to 71.5), grade C+, 18 comparable benchmarks, verification: third-party published. Reasoning 80.0, Knowledge 87.6 (single-test category, shown beside the Index, not averaged), Coding 69.3, Math 55.3, Agentic 76.6 (not counted: fewer than half the category's tests), Multimodal 91.7 (single-test category, shown beside the Index, not averaged), Human preference 92.0 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gemini-3-5-flash #19 (17 to 25). GPT-5.2 (OpenAI, United States) — Index 67.7 (leave-one-out 63.0 to 72.0), grade C+, 21 comparable benchmarks, verification: third-party published. Reasoning 70.5, Knowledge 49.1 (single-test category, shown beside the Index, not averaged), Coding 62.1 (not counted: fewer than half the category's tests), Math 53.9, Agentic 78.6, Multimodal 80.4 (single-test category, shown beside the Index, not averaged), Human preference 80.3 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gpt-5-2 #20 (17 to 23). DeepSeek V4 Pro (DeepSeek, China) — Index 67.5 (leave-one-out 63.3 to 71.2), grade C+, 16 comparable benchmarks, verification: third-party published. Reasoning 76.6, Knowledge 70.0 (single-test category, shown beside the Index, not averaged), Coding 64.6, Math 61.3, Agentic 60.6 (not counted: fewer than half the category's tests), Human preference 86.4 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/deepseek-v4-pro #21 (21 to 28). Grok 4.5 (xAI, United States) — Index 64.2 (leave-one-out 59.3 to 67.4), grade C, 18 comparable benchmarks, verification: third-party published. Reasoning 74.8, Knowledge 63.9 (single-test category, shown beside the Index, not averaged), Coding 64.1, Math 53.7, Agentic 69.1 (not counted: fewer than half the category's tests), Multimodal 91.1 (single-test category, shown beside the Index, not averaged), Human preference 89.7 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/grok-4-5 #22 (18 to 28). Claude Sonnet 4.6 (Anthropic, United States) — Index 63.9 (leave-one-out 61.0 to 69.7), grade C, 20 comparable benchmarks, verification: third-party published. Reasoning 62.3, Knowledge 47.0 (single-test category, shown beside the Index, not averaged), Coding 63.8, Math 65.4 (not counted: fewer than half the category's tests), Agentic 65.6, Multimodal 89.2 (single-test category, shown beside the Index, not averaged), Human preference 90.0 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-sonnet-4-6 #23 (20 to 28). DeepSeek V4 Flash (DeepSeek, China) — Index 63.7 (leave-one-out 59.8 to 67.5), grade C, 15 comparable benchmarks, verification: third-party published. Reasoning 74.8, Knowledge 44.5 (single-test category, shown beside the Index, not averaged), Coding 57.0, Math 59.2, Human preference 80.5 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/deepseek-v4-flash #24 (21 to 28). GLM 5.2 (Zhipu AI, China) — Index 63.3 (leave-one-out 59.6 to 66.9), grade C, 19 comparable benchmarks, verification: third-party published. Reasoning 66.6, Knowledge 45.2 (single-test category, shown beside the Index, not averaged), Coding 69.6, Math 53.7, Agentic 82.8 (not counted: fewer than half the category's tests), Human preference 89.8 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/glm-5-2 #25 (20 to 28). Gemini 3.6 Flash (Google, United States) — Index 63.2 (leave-one-out 58.7 to 67.5), grade C, 16 comparable benchmarks, verification: third-party published. Reasoning 71.7, Knowledge 87.6 (single-test category, shown beside the Index, not averaged), Coding 63.9, Math 53.9, Multimodal 91.9 (single-test category, shown beside the Index, not averaged), Human preference 92.3 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gemini-3-6-flash #26 (21 to 30). Gemini 3 Flash (Google, United States) — Index 62.5 (leave-one-out 57.0 to 66.6), grade C, 18 comparable benchmarks, verification: third-party published. Reasoning 71.8, Knowledge 88.4 (single-test category, shown beside the Index, not averaged), Coding 59.7 (not counted: fewer than half the category's tests), Math 45.7, Agentic 70.1, Multimodal 88.4 (single-test category, shown beside the Index, not averaged), Human preference 90.4 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gemini-3-flash #27 (21 to 30). Claude Opus 4.5 (Anthropic, United States) — Index 62.4 (leave-one-out 57.4 to 66.4), grade C, 21 comparable benchmarks, verification: third-party published. Reasoning 67.4, Knowledge 60.4 (single-test category, shown beside the Index, not averaged), Coding 63.4 (not counted: fewer than half the category's tests), Math 41.0, Agentic 78.9, Human preference 90.2 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-opus-4-5 #28 (21 to 28). GPT-5 (OpenAI, United States) — Index 62.0 (leave-one-out 59.4 to 65.0), grade C, 25 comparable benchmarks, verification: third-party published. Reasoning 54.4, Knowledge 66.3 (single-test category, shown beside the Index, not averaged), Coding 68.6, Math 58.2, Agentic 66.7, Multimodal 71.3 (single-test category, shown beside the Index, not averaged), Human preference 79.4 (single-test category, shown beside the Index, not averaged), Long context 96.9 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gpt-5 #29 (29 to 30). Kimi K2.7 Code (Moonshot AI, China) — Index 58.6 (leave-one-out 56.2 to 61.2), grade C-, 14 comparable benchmarks, verification: third-party published. Reasoning 63.5, Knowledge 48.3 (single-test category, shown beside the Index, not averaged), Coding 57.1, Math 55.3, Agentic 66.4 (not counted: fewer than half the category's tests). https://opencharts.com/benchmarks/kimi-k2-7-code #30 (29 to 30). GLM 5.1 (Zhipu AI, China) — Index 57.5 (leave-one-out 54.7 to 60.5), grade C-, 13 comparable benchmarks, verification: third-party published. Reasoning 57.7, Knowledge 45.0 (single-test category, shown beside the Index, not averaged), Coding 63.2, Math 51.6, Agentic 77.9 (not counted: fewer than half the category's tests), Human preference 88.2 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/glm-5-1 #31 (31 to 32). Claude Sonnet 4.5 (Anthropic, United States) — Index 54.3 (leave-one-out 50.9 to 57.0), grade C-, 23 comparable benchmarks, verification: third-party published. Reasoning 43.2, Knowledge 40.6 (single-test category, shown beside the Index, not averaged), Coding 56.6, Math 44.9, Agentic 72.5, Human preference 85.5 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-sonnet-4-5 #32 (31 to 32). GPT-5.1 (OpenAI, United States) — Index 54.0 (leave-one-out 51.5 to 56.4), grade C-, 18 comparable benchmarks, verification: third-party published. Reasoning 52.0, Knowledge 63.5 (single-test category, shown beside the Index, not averaged), Coding 58.6, Math 88.6 (not counted: fewer than half the category's tests), Agentic 51.3, Multimodal 82.0 (single-test category, shown beside the Index, not averaged), Human preference 85.1 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gpt-5-1 #33 (28 to 34). Claude Opus 4.1 (Anthropic, United States) — Index 50.7 (leave-one-out 47.4 to 62.0), grade C-, 15 comparable benchmarks, verification: third-party published. Reasoning 57.3, Coding 51.6 (not counted: fewer than half the category's tests), Math 28.3, Agentic 66.6, Human preference 83.6 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/claude-opus-4-1 #34 (33 to 34). GPT-5.4 mini (OpenAI, United States) — Index 49.3 (leave-one-out 44.3 to 53.0), grade D, 17 comparable benchmarks, verification: third-party published. Reasoning 50.7, Knowledge 38.9 (single-test category, shown beside the Index, not averaged), Coding 53.5, Math 43.7, Agentic 58.8 (not counted: fewer than half the category's tests), Multimodal 82.6 (single-test category, shown beside the Index, not averaged), Human preference 83.2 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gpt-5-4-mini #35 (35 to 36). Inkling (Thinking Machines, United States) — Index 42.8 (leave-one-out 36.5 to 46.4), grade D, 15 comparable benchmarks, verification: third-party published. Reasoning 57.3, Knowledge 53.3 (single-test category, shown beside the Index, not averaged), Coding 38.7, Math 32.4, Human preference 80.7 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/inkling #36. Gemini 2.5 Pro (Google, United States) — Index 39.5 (leave-one-out 35.2 to 41.9), grade F, 18 comparable benchmarks, verification: third-party published. Reasoning 34.5, Coding 50.8, Math 36.8, Agentic 35.7, Multimodal 81.1 (single-test category, shown beside the Index, not averaged), Human preference 82.5 (single-test category, shown beside the Index, not averaged). https://opencharts.com/benchmarks/gemini-2-5-pro ## Provisional (too few benchmarks to rank) - GPT-5.5 Pro (OpenAI) — Index 93.9 from 6 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-5-pro - GPT-5.4 Pro (OpenAI) — Index 93.4 from 10 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-4-pro - ERNIE 5.1 (Baidu) — Index 88.8 from 1 benchmark, verification: third-party published. https://opencharts.com/benchmarks/ernie-5-1 - Gemini 3 Deep Think (Google) — Index 88.7 from 3 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gemini-3-deep-think - GPT-5.6 Sol Pro (OpenAI) — Index 87.7 from 4 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-6-sol-pro - Seed 2.0 Pro (ByteDance) — Index 84.7 from 2 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/seed-2-0-pro - GPT-5.3 Codex (OpenAI) — Index 80.9 from 7 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-3-codex - Qwen3.6 Max (Alibaba) — Index 76.3 from 8 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-6-max - Muse Spark 1.2 (Meta) — Index 75.2 from 9 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/muse-spark-1-2 - GPT-5.2 Pro (OpenAI) — Index 73.0 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-2-pro - Gemini 3 Pro (Google) — Index 72.3 from 19 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gemini-3-pro - Muse Spark (Meta) — Index 71.3 from 8 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/muse-spark - Muse Spark 1.1 (Meta) — Index 71.3 from 9 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/muse-spark-1-1 - Qwen3.8 Max (Alibaba) — Index 71.2 from 11 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-8-max - Gemini 3.8 Flash (Google) — Index 68.9 from 6 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gemini-3-8-flash - Qwen3.7 Flash (Alibaba) — Index 67.4 from 3 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-7-flash - Qwen3.5 397B-A17B (Alibaba) — Index 65.7 from 6 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-5-397b-a17b - Qwen3.7 Max (Alibaba) — Index 65.0 from 12 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-7-max - Qwen3 Max (Alibaba) — Index 64.2 from 7 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-max - Qwen3.6 Plus (Alibaba) — Index 64.1 from 10 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-6-plus - Grok 4.20 (xAI) — Index 63.2 from 14 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/grok-4-20 - Gemma 4 26B A4B (Google) — Index 63.1 from 8 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gemma-4-26b-a4b - GLM 5.3 (Zhipu AI) — Index 60.9 from 12 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/glm-5-3 - Kimi K2 (Moonshot AI) — Index 60.6 from 10 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/kimi-k2 - Hunyuan 3 (Tencent) — Index 58.8 from 2 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/hunyuan-3 - GPT-5 Pro (OpenAI) — Index 58.5 from 7 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-pro - Kimi K2.6 (Moonshot AI) — Index 58.5 from 17 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/kimi-k2-6 - Qwen3.7 Plus (Alibaba) — Index 58.4 from 9 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-7-plus - Gemma 4 31B (Google) — Index 57.0 from 9 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gemma-4-31b - Grok 4 (xAI) — Index 56.3 from 16 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/grok-4 - Qwen3.6 27B (Alibaba) — Index 56.3 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-6-27b - Qwen3.5 Plus (Alibaba) — Index 54.8 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-5-plus - Qwen3.8 Flash Next (Alibaba) — Index 54.5 from 1 benchmark, verification: third-party published. https://opencharts.com/benchmarks/qwen3-8-flash-next - Muse Spark 1.3 (Meta) — Index 53.6 from 1 benchmark, verification: third-party published. https://opencharts.com/benchmarks/muse-spark-1-3 - MiniMax M3 (MiniMax) — Index 53.5 from 12 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/minimax-m3 - GPT-5 Codex (OpenAI) — Index 53.0 from 3 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-codex - Qwen3.5 35B-A3B (Alibaba) — Index 52.2 from 6 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-5-35b-a3b - Kimi K2.5 (Moonshot AI) — Index 52.1 from 19 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/kimi-k2-5 - GLM 5 (Zhipu AI) — Index 51.1 from 12 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/glm-5 - Qwen3.8 27B (Alibaba) — Index 51.0 from 6 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-8-27b - Inkling Small (Thinking Machines) — Index 50.8 from 13 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/inkling-small - Nemotron 3 Ultra (NVIDIA) — Index 49.8 from 8 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/nemotron-3-ultra - GLM 4.7 (Zhipu AI) — Index 48.6 from 12 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/glm-4-7 - DeepSeek V3.1 Terminus (DeepSeek) — Index 48.5 from 3 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/deepseek-v3-1-terminus - DeepSeek V3.2 (DeepSeek) — Index 48.3 from 15 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/deepseek-v3-2 - GLM 5.3 Flash (Zhipu AI) — Index 48.0 from 10 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/glm-5-3-flash - Qwen3.5 122B-A10B (Alibaba) — Index 47.3 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-5-122b-a10b - Grok 4 Fast (xAI) — Index 47.1 from 7 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/grok-4-fast - MiMo V2.5 (Xiaomi) — Index 46.1 from 6 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/mimo-v2-5 - Qwen3.6 Flash (Alibaba) — Index 46.0 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-6-flash - Gemini 2.5 Flash (Google) — Index 45.2 from 8 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gemini-2-5-flash - MiMo V2.5 Pro (Xiaomi) — Index 44.3 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/mimo-v2-5-pro - Mistral Large 3 (Mistral AI) — Index 44.1 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/mistral-large-3 - Qwen3.5 27B (Alibaba) — Index 43.6 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-5-27b - Grok 4.1 (xAI) — Index 43.1 from 4 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/grok-4-1 - GPT-5 mini (OpenAI) — Index 43.0 from 19 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-mini - Qwen3.6 35B-A3B (Alibaba) — Index 42.9 from 7 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-6-35b-a3b - MiniMax M2.1 (MiniMax) — Index 42.1 from 3 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/minimax-m2-1 - Grok 4.1 Fast (xAI) — Index 41.9 from 6 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/grok-4-1-fast - Grok 4.3 (xAI) — Index 41.3 from 13 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/grok-4-3 - GPT-5.2 Codex (OpenAI) — Index 40.9 from 3 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-2-codex - GPT-5.1 Codex Mini (OpenAI) — Index 40.3 from 2 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-1-codex-mini - GPT-5.4 nano (OpenAI) — Index 39.5 from 14 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-4-nano - Qwen3.5 Flash (Alibaba) — Index 39.5 from 7 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-5-flash - GPT-OSS 20B (OpenAI) — Index 39.4 from 7 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-oss-20b - Qwen3.5 9B (Alibaba) — Index 38.9 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/qwen3-5-9b - GPT-5 nano (OpenAI) — Index 35.8 from 14 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-nano - MiniMax M2.7 (MiniMax) — Index 35.4 from 7 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/minimax-m2-7 - GLM 4.6 (Zhipu AI) — Index 34.2 from 6 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/glm-4-6 - Gemini 3.1 Flash-Lite (Google) — Index 34.1 from 12 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gemini-3-1-flash-lite - Mistral Medium 3.5 (Mistral AI) — Index 33.7 from 8 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/mistral-medium-3-5 - GPT-5.5 Instant (OpenAI) — Index 32.9 from 8 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-5-instant - Claude Haiku 4.5 (Anthropic) — Index 32.7 from 15 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/claude-haiku-4-5 - Gemini 3.5 Flash-Lite (Google) — Index 32.5 from 13 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gemini-3-5-flash-lite - Nemotron 3 Super (NVIDIA) — Index 29.6 from 3 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/nemotron-3-super - MiniMax M2.5 (MiniMax) — Index 28.9 from 8 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/minimax-m2-5 - GPT-5.1 Codex (OpenAI) — Index 28.8 from 5 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-5-1-codex - MiniMax M2 (MiniMax) — Index 28.3 from 4 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/minimax-m2 - GPT-OSS 120B (OpenAI) — Index 27.8 from 12 benchmarks, verification: third-party published. https://opencharts.com/benchmarks/gpt-oss-120b ## Listed without scores yet - TheoArca Magnus 5.1 (TheoVex) — first-party engine; independent verification TBD; internal evaluation pending. https://opencharts.com/benchmarks/theoarca-magnus-5-1 - TheoArca Velox 5.1 (TheoVex) — first-party engine; independent verification TBD; internal evaluation pending. https://opencharts.com/benchmarks/theoarca-velox-5-1 ## Category leaders (third-party published) - Reasoning: GPT-6 Astra (OpenAI) — 99.6 - Knowledge: GPT-6 Astra (OpenAI) — 100.0 (single-test category, beside the Index) - Coding: GPT-6 Astra (OpenAI) — 97.7 - Math: GPT-6 Astra (OpenAI) — 99.8 - Agentic: Claude Opus 4.6 (Anthropic) — 91.1 - Multimodal: Claude Fable 5 (Anthropic) — 100.0 (single-test category, beside the Index) - Human preference: Claude Fable 5 (Anthropic) — 100.0 (single-test category, beside the Index) - Long context: GPT-5 (OpenAI) — 96.9 (single-test category, beside the Index) ## Hard set (mean share of the frontier on the hardest tests) 1. Claude Fable 5.1 (Anthropic) — hard-set score 95.3 across 4 of 8 tests 2. GPT-5.5 (OpenAI) — hard-set score 88.7 across 4 of 8 tests 3. GPT-5.4 Pro (OpenAI) — hard-set score 82.8 across 4 of 8 tests 4. GPT-5.4 (OpenAI) — hard-set score 80.3 across 7 of 8 tests 5. Gemini 3.1 Pro (Google) — hard-set score 77.6 across 7 of 8 tests 6. Claude Opus 4.7 (Anthropic) — hard-set score 76.6 across 6 of 8 tests 7. Claude Opus 4.6 (Anthropic) — hard-set score 75.4 across 7 of 8 tests 8. Gemini 3 Pro (Google) — hard-set score 73.6 across 6 of 8 tests 9. GPT-5.2 (OpenAI) — hard-set score 71.2 across 8 of 8 tests 10. Gemini 3.5 Flash (Google) — hard-set score 66.4 across 4 of 8 tests Waiting on more results (fewer than 4 hard-set tests published so far): GPT-6 Astra 100.0 on 3; Claude Fable 5 93.1 on 3; GPT-5.6 Sol 92.5 on 3; GPT-5.3 Codex 88.8 on 3; GPT-5.5 Pro 87.5 on 3; Claude Opus 5 87.2 on 3. ## Per-model scores ### GPT-6 Astra (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-6-astra Index: 99.0 (leave-one-out 99.0 to 99.7) · Grade: A+ · Rank: #1 · Verification: third-party published - GPQA Diamond: 95.8% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 95.0% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 98.5% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 31.7% (98.2: 98.2% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 75.6% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SciCode: 56.5% (91.0: 91.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 92.9% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 53.3% (99.6: 99.6% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Code Arena (WebDev): 1,797 (100.0: best published result) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 100.0% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 93.7% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 97.6% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - ProofBench: 99.0% (99.0: 99.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 46.7% (98.5: 98.5% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents ### Claude Fable 5.1 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-fable-5-1 Index: 97.0 (leave-one-out 96.6 to 97.8) · Grade: A+ · Rank: #2 · Verification: third-party published - Humanity's Last Exam: 46.5% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 90.0% (94.7: 94.7% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 97.5% (99.0: 99.0% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 31.1% (96.4: 96.4% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 70.8% (93.7: 93.7% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SciCode: 62.0% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 92.9% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 50.9% (95.2: 95.2% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 73.4% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,762 (90.0: 45% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 100.0% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 90.2% (96.3: 96.3% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 87.8% (90.0: 90.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - ProofBench: 100.0% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 47.4% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $5,422 (76.7: 76.7% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,504 (99.2: 50% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Fable 5 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-fable-5 Index: 92.7 (leave-one-out 91.9 to 95.6) · Grade: A · Rank: #3 (3 to 4) · Verification: third-party published - GPQA Diamond: 85.9% (86.0: 86.0% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FS5pm7ctk9zgPfbyHrbD9AV.eval - ARC-AGI-2: 89.2% (93.9: 93.9% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 98.5% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 81.9% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 28.6% (88.5: 88.5% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 70.7% (93.5: 93.5% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SciCode: 60.2% (97.0: 97.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 91.9% (99.0: 99.0% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 53.5% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 70.5% (96.0: 96.0% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,629 (55.1: 28% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 100.0% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FhwrnvtEbDAocFtjcsHqsQh.eval - FrontierMath (Tiers 1–3): 87.0% (92.9: 92.9% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FE28uZiJvrPgKtawbJpoZQ8.eval - FrontierMath Tier 4: 90.2% (92.4: 92.4% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FReDAtJGGfRtrEGGWf4Zbsd.eval - ProofBench: 95.0% (95.0: 95.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 45.0% (94.9: 94.9% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $5,680 (78.2: 78.2% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,313 (100.0: best published result) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,507 (100.0: best published result) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Opus 5 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-opus-5 Index: 92.6 (leave-one-out 91.7 to 94.3) · Grade: A · Rank: #4 (3 to 4) · Verification: third-party published - GPQA Diamond: 93.9% (97.3: 97.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 90.4% (95.2: 95.2% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 97.5% (99.0: 99.0% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 80.6% (98.4: 98.4% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 29.1% (90.2: 90.2% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 59.9% (79.2: 79.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SciCode: 55.7% (89.7: 89.7% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 91.8% (98.8: 98.8% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 53.4% (99.8: 99.8% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 70.0% (95.4: 95.4% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,688 (69.5: 35% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 98.9% (98.9: 98.9% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 85.6% (91.4: 91.4% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fff7GeS7rdperjxSBZ7MHZY.eval - FrontierMath Tier 4: 73.2% (75.0: 75.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fji3kMn49Jcr9t4whxcpjj4.eval - ProofBench: 99.0% (99.0: 99.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - OSWorld 2.0: 31.4% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://osworld-v2.xlang.ai/ - APEX-Agents: 43.5% (91.8: 91.8% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $11,182 (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,290 (93.5: 47% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,493 (96.0: 48% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.6 Sol (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-6-sol Index: 89.8 (leave-one-out 88.7 to 92.4) · Grade: A- · Rank: #5 · Verification: third-party published - GPQA Diamond: 93.5% (96.8: 96.8% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 92.5% (97.4: 97.4% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 97.5% (99.0: 99.0% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 64.8% (79.1: 79.1% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 32.3% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 69.7% (92.2: 92.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SciCode: 56.9% (91.8: 91.8% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 88.8% (95.5: 95.5% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 47.5% (88.8: 88.8% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 67.2% (91.6: 91.6% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,617 (52.5: 26% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 100.0% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 89.1% (95.1: 95.1% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FVBpT2WWYKBZtUxG9Z3svGh.eval - FrontierMath Tier 4: 82.9% (85.0: 85.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FQJ5rJMVypB4PMxPcBGftc8.eval - ProofBench: 83.0% (83.0: 83.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - OSWorld 2.0: 27.3% (87.0: 87.0% of the best (31.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://osworld-v2.xlang.ai/ - APEX-Agents: 39.9% (84.2: 84.2% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $9,619 (95.2: 95.2% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,282 (91.2: 46% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,483 (93.1: 47% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.6 Terra (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-6-terra Index: 81.9 (leave-one-out 80.2 to 85.3) · Grade: B+ · Rank: #6 · Verification: third-party published - GPQA Diamond: 93.3% (96.5: 96.5% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 83.9% (88.3: 88.3% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 96.5% (98.0: 98.0% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 48.9% (59.7: 59.7% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 30.0% (92.9: 92.9% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 43.2% (57.1: 57.1% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SciCode: 53.9% (86.9: 86.9% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 78.3% (84.3: 84.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 41.3% (77.2: 77.2% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 64.9% (88.4: 88.4% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,520 (33.7: 17% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 99.7% (99.7: 99.7% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 86.0% (91.8: 91.8% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FJ6evjTMgHkrun8txnEvzdN.eval - FrontierMath Tier 4: 70.7% (72.5: 72.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FZUygpkXqg3NjchXXfedAWV.eval - ProofBench: 74.0% (74.0: 74.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vending-Bench 2: $7,343 (86.5: 86.5% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,266 (86.5: 43% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,466 (88.3: 44% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.5 (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-5 Index: 80.2 (leave-one-out 78.8 to 82.8) · Grade: B+ · Rank: #7 (6 to 9) · Verification: third-party published - GPQA Diamond: 90.7% (92.8: 92.8% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F2CdomqtBVz2it7a2Zx2G2t.eval - ARC-AGI-2: 85.0% (89.5: 89.5% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 95.0% (96.4: 96.4% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 69.0% (84.2: 84.2% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lmcouncil.ai/benchmarks - CritPt: 27.1% (84.0: 84.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 63.0% (83.3: 83.3% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FTijCafTTJUFrJ6jPGvBbgq.eval - SciCode: 56.1% (90.5: 90.5% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 84.9% (91.4: 91.4% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 43.0% (80.3: 80.3% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 58.4% (79.6: 79.6% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,510 (32.1: 16% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 84.4% (84.4: 84.4% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FCdmtwZ8NR8YeTXWiLVTRDJ.eval - FrontierMath (Tiers 1–3): 85.3% (91.0: 91.0% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 72.5% (74.3: 74.3% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - ProofBench: 50.0% (50.0: 50.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 84.7% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - OSWorld 2.0: 13.0% (41.4: 41.4% of the best (31.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://osworld-v2.xlang.ai/ - APEX-Agents: 38.5% (81.2: 81.2% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $7,524 (87.2: 87.2% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 54.0% (97.6: 97.6% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,286 (92.3: 46% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,482 (92.8: 46% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Kimi K3 (Moonshot AI, China) URL: https://opencharts.com/benchmarks/kimi-k3 Index: 79.6 (leave-one-out 77.2 to 83.5) · Grade: B · Rank: #8 (6 to 10) · Verification: third-party published - GPQA Diamond: 93.1% (96.3: 96.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 60.4% (63.6: 63.6% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 94.5% (95.9: 95.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 60.7% (74.1: 74.1% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 23.4% (72.4: 72.4% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 50.6% (66.9: 66.9% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmZLWt2JSYvJyGhMvji7CT8.eval - SciCode: 58.7% (94.6: 94.6% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 82.6% (88.9: 88.9% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 44.2% (82.6: 82.6% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 60.8% (82.8: 82.8% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,674 (66.1: 33% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 97.2% (97.2: 97.2% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 72.2% (77.0: 77.0% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 39.0% (40.0: 40.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FgZERSAQwxBAevQzBsWfUie.eval - ProofBench: 87.0% (87.0: 87.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 39.3% (82.9: 82.9% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $5,165 (75.1: 75.1% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,489 (94.7: 47% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Opus 4.8 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-opus-4-8 Index: 79.5 (leave-one-out 77.7 to 81.8) · Grade: B · Rank: #9 (7 to 9) · Verification: third-party published - GPQA Diamond: 91.0% (93.3: 93.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 72.1% (75.9: 75.9% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 92.5% (93.9: 93.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 64.8% (79.1: 79.1% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 20.9% (64.6: 64.6% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 53.0% (70.1: 70.1% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FViC6pxQZUVS3E9M9iEubF7.eval - SciCode: 53.5% (86.2: 86.2% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 82.9% (89.2: 89.2% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 46.5% (86.9: 86.9% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 62.3% (84.9: 84.9% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,563 (41.2: 21% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 98.3% (98.3: 98.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 80.0% (85.4: 85.4% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FGVmgtgKMwEqc54f3jA7auP.eval - FrontierMath Tier 4: 56.1% (57.5: 57.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FJiqgqBTfkGx797uowZgJbk.eval - ProofBench: 69.0% (69.0: 69.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - OSWorld 2.0: 20.6% (65.5: 65.5% of the best (31.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://osworld-v2.xlang.ai/ - APEX-Agents: 42.5% (89.7: 89.7% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $5,787 (78.8: 78.8% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 50.2% (90.8: 90.8% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,285 (92.0: 46% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,482 (92.9: 46% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.4 (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-4 Index: 77.3 (leave-one-out 75.2 to 79.1) · Grade: B · Rank: #10 (10 to 11) · Verification: third-party published - GPQA Diamond: 93.3% (96.5: 96.5% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FfFatyce8UvpN7ZivdmrhAy.eval - Humanity's Last Exam: 36.2% (77.9: 77.9% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 74.0% (77.8: 77.8% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 93.7% (95.1: 95.1% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 23.4% (72.5: 72.5% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 45.1% (59.7: 59.7% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FAzCkm282PLhDdLJ8DnAQQc.eval - SWE-bench Verified: 76.9% (92.1: 92.1% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fi9HTMbAxcwsrTf2w2FuJ6N.eval - SciCode: 56.6% (91.2: 91.2% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 77.7% (83.6: 83.6% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,463 (25.5: 13% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 97.8% (97.8: 97.8% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FDpPahiEawokzgidE3cNZbS.eval - FrontierMath (Tiers 1–3): 78.6% (83.9: 83.9% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FPaYEQ37A2n6giYBwZGBnXY.eval - FrontierMath Tier 4: 49.0% (50.2: 50.2% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FdoiQDghqnLzbE8CFijLNQJ.eval - ProofBench: 56.0% (56.0: 56.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 81.8% (96.6: 96.6% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - METR Time Horizon: 5.7 h (83.9: 83.9% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/research - APEX-Agents: 36.0% (75.9: 75.9% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $6,144 (80.7: 80.7% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 35.1% (63.5: 63.5% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,281 (90.7: 45% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,477 (91.2: 46% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Opus 4.6 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-opus-4-6 Index: 75.4 (leave-one-out 72.6 to 78.4) · Grade: B · Rank: #11 (10 to 16) · Verification: third-party published - GPQA Diamond: 90.5% (92.6: 92.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FDiXCMhdfhcjaCUSzYWgACy.eval - Humanity's Last Exam: 34.4% (74.1: 74.1% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 69.2% (72.8: 72.8% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 94.0% (95.4: 95.4% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 67.6% (82.5: 82.5% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - SimpleQA Verified: 47.0% (62.2: 62.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F85fhMxLaebWtLoi6oLpTbi.eval - SWE-bench Verified: 78.7% (94.3: 94.3% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F2hH2GAPUFaJR7pfVKhhicg.eval - WeirdML: 78.0% (83.9: 83.9% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 26.6% (49.8: 49.8% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Code Arena (WebDev): 1,546 (38.2: 19% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 94.4% (94.4: 94.4% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FTUPC8VdX4jApFa4MGfNP4t.eval - FrontierMath (Tiers 1–3): 66.0% (70.4: 70.4% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FkgJ2d6XbKj3dZutDcBYSKU.eval - FrontierMath Tier 4: 26.8% (27.5: 27.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F2hbbp8z6PhB7c2T33os9q5.eval - ProofBench: 50.0% (50.0: 50.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 79.8% (94.2: 94.2% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - METR Time Horizon: 12.0 h (94.6: 94.6% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/research - APEX-Agents: 32.4% (68.4: 68.4% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $8,018 (89.3: 89.3% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Cybench: 93.0% (100.0: best published result) — developer-reported; source: Epoch AI — AI Benchmarking Hub, https://www-cdn.anthropic.com/0dd865075ad3132672ee0ab40b05a53f14cf5288.pdf - DeepResearch Bench: 55.3% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,299 (95.9: 48% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,505 (99.3: 50% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 3.7 Flash (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-7-flash Index: 75.2 (leave-one-out 71.9 to 79.5) · Grade: B · Rank: #12 (9 to 16) · Verification: third-party published - GPQA Diamond: 94.8% (98.7: 98.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 84.6% (89.0: 89.0% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 95.5% (97.0: 97.0% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 14.3% (44.2: 44.2% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 69.2% (91.5: 91.5% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F84ARuGyhQsjgh8umRmBvas.eval - SciCode: 57.9% (93.3: 93.3% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - FrontierCode: 43.6% (81.5: 81.5% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 61.6% (83.9: 83.9% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,587 (46.0: 23% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 97.2% (97.2: 97.2% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 71.6% (76.4: 76.4% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FaxMdgNDC8UYYLnZYLHZpF5.eval - FrontierMath Tier 4: 36.6% (37.5: 37.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FKptKXK5nE2Ns2WbiaSSnbc.eval - ProofBench: 58.0% (58.0: 58.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Text Arena: 1,491 (95.3: 48% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Grok 4.6 (xAI, United States) URL: https://opencharts.com/benchmarks/grok-4-6 Index: 75.2 (leave-one-out 71.2 to 78.6) · Grade: B · Rank: #13 (10 to 16) · Verification: third-party published - GPQA Diamond: 94.0% (97.5: 97.5% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 67.1% (70.6: 70.6% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 87.5% (88.8: 88.8% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 75.9% (92.7: 92.7% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 19.7% (61.0: 61.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 49.3% (65.2: 65.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FTpHYtuPG6kYZFvMK7Z6Hqj.eval - SciCode: 54.6% (88.1: 88.1% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 67.3% (72.4: 72.4% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 48.0% (89.8: 89.8% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 70.8% (96.5: 96.5% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,625 (54.1: 27% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 99.2% (99.2: 99.2% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 66.0% (70.4: 70.4% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FZb2AKwa942D7dBTTcebAba.eval - FrontierMath Tier 4: 31.7% (32.5: 32.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fbg3JRdd33iWfvRoGSvyXXE.eval - ProofBench: 51.0% (51.0: 51.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 41.2% (86.9: 86.9% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $9,047 (93.2: 93.2% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 ### Claude Opus 4.7 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-opus-4-7 Index: 75.1 (leave-one-out 72.4 to 77.8) · Grade: B · Rank: #14 (10 to 16) · Verification: third-party published - GPQA Diamond: 90.2% (92.1: 92.1% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FVJezHrJ5YPCbpui8pZgcJg.eval - Humanity's Last Exam: 36.2% (77.8: 77.8% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 75.8% (79.8: 79.8% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 93.5% (94.9: 94.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 61.7% (75.3: 75.3% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lmcouncil.ai/benchmarks - CritPt: 12.0% (37.2: 37.2% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 51.7% (68.4: 68.4% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F2c4rvmf5N7QXTugrFLMZ83.eval - SWE-bench Verified: 83.5% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FnCmGWWMBip2s9AVEZNYXDX.eval - SciCode: 54.5% (87.9: 87.9% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 76.4% (82.3: 82.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 38.5% (72.1: 72.1% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 64.8% (88.3: 88.3% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,557 (40.2: 20% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 97.8% (97.8: 97.8% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjyCZnpXG2bh9G43X8pT75m.eval - FrontierMath (Tiers 1–3): 70.2% (74.9: 74.9% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FZkNQ7GER7Mc9FSPmpUzgKH.eval - FrontierMath Tier 4: 31.7% (32.5: 32.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FP4LXQ9PyDDziEANMB5jVRd.eval - ProofBench: 54.0% (54.0: 54.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 80.2% (94.7: 94.7% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - OSWorld 2.0: 18.2% (57.9: 57.9% of the best (31.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://osworld-v2.xlang.ai/ - APEX-Agents: 33.9% (71.5: 71.5% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $10,937 (99.3: 99.3% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,301 (96.7: 48% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,502 (98.5: 49% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 3.1 Pro (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-1-pro Index: 73.4 (leave-one-out 69.9 to 75.7) · Grade: B- · Rank: #15 (11 to 17) · Verification: third-party published - GPQA Diamond: 94.4% (98.1: 98.1% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FBgmrBFsbsaD8rNpebv49Da.eval - Humanity's Last Exam: 46.4% (99.9: 99.9% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 77.1% (81.2: 81.2% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 98.0% (99.5: 99.5% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 79.6% (97.2: 97.2% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - CritPt: 17.7% (54.8: 54.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 73.5% (97.2: 97.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SWE-bench Verified: 75.6% (90.6: 90.6% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F8QQQWDgmmEsmQVUJWcxx4P.eval - SciCode: 58.9% (95.0: 95.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 72.1% (77.6: 77.6% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,446 (23.4: 12% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 95.6% (95.6: 95.6% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FbXgZZgzRE3wXoSRUgrbsVf.eval - FrontierMath (Tiers 1–3): 59.6% (63.7: 63.7% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FhVifTUdQZrFEYabgCRu8Qq.eval - FrontierMath Tier 4: 26.8% (27.5: 27.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FeXVWr3RBeSsWy5DAU8ed46.eval - ProofBench: 26.0% (26.0: 26.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 80.2% (94.7: 94.7% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - METR Time Horizon: 6.4 h (85.6: 85.6% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/research - APEX-Agents: 33.5% (70.7: 70.7% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $3,774 (65.0: 65.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 47.8% (86.4: 86.4% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,278 (90.1: 45% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,487 (94.1: 47% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.6 Luna (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-6-luna Index: 72.9 (leave-one-out 70.6 to 75.8) · Grade: B- · Rank: #16 (11 to 17) · Verification: third-party published - GPQA Diamond: 91.6% (94.1: 94.1% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 59.5% (62.7: 62.7% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 88.0% (89.3: 89.3% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 46.8% (57.1: 57.1% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 20.6% (63.8: 63.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 41.0% (54.2: 54.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SciCode: 52.5% (84.7: 84.7% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 60.9% (65.5: 65.5% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 39.8% (74.4: 74.4% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 61.1% (83.2: 83.2% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,519 (33.7: 17% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 98.3% (98.3: 98.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 82.1% (87.6: 87.6% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FYLAk5J3Wd8EruVXee8bAQb.eval - FrontierMath Tier 4: 61.0% (62.5: 62.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FNLiiJJSyA9wmg2FHENFpAq.eval - ProofBench: 60.0% (60.0: 60.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vending-Bench 2: $4,095 (67.7: 67.7% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,254 (83.0: 42% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,453 (84.4: 42% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Sonnet 5 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-sonnet-5 Index: 71.0 (leave-one-out 68.0 to 75.2) · Grade: B- · Rank: #17 (12 to 18) · Verification: third-party published - GPQA Diamond: 90.5% (92.6: 92.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - SimpleBench: 60.6% (74.0: 74.0% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 16.9% (52.2: 52.2% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 33.7% (44.6: 44.6% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FLPo74T9L3CawnKX5LLUrSZ.eval - SciCode: 53.6% (86.4: 86.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 68.8% (74.0: 74.0% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 42.7% (79.9: 79.9% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 61.5% (83.8: 83.8% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,537 (36.6: 18% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 94.7% (94.7: 94.7% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 65.6% (70.0: 70.0% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FV6YuGpcu8MCbJNXKoJMA6b.eval - FrontierMath Tier 4: 29.3% (30.0: 30.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FbqiVinHXFcgyFGPstJwXJN.eval - ProofBench: 77.0% (77.0: 77.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 32.5% (68.6: 68.6% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $6,378 (81.9: 81.9% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,267 (87.0: 44% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,462 (87.2: 44% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 3.5 Flash (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-5-flash Index: 68.2 (leave-one-out 63.7 to 71.5) · Grade: C+ · Rank: #18 (17 to 22) · Verification: third-party published - GPQA Diamond: 92.8% (95.8: 95.8% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FdYYtcwENFRtm9jZTDA2p9C.eval - ARC-AGI-2: 72.1% (75.9: 75.9% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 92.5% (93.9: 93.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 76.7% (93.7: 93.7% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 13.1% (40.7: 40.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 66.2% (87.6: 87.6% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FGJ6YsSuzyRD9YXe6SBwvsu.eval - SWE-bench Verified: 79.3% (95.0: 95.0% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FdwRzFGTghYEy4HNZajJXoc.eval - SciCode: 53.1% (85.6: 85.6% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 62.6% (67.4: 67.4% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - CursorBench: 49.8% (67.8: 67.8% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,500 (30.7: 15% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 95.6% (95.6: 95.6% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FeYmgjYJfM7Utm46nfoRfZ6.eval - FrontierMath (Tiers 1–3): 62.8% (67.0: 67.0% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F72DQVwbYiQW5pjhshjBaGU.eval - FrontierMath Tier 4: 26.8% (27.5: 27.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FcXn5zzns6VFwbTgotDwZjV.eval - ProofBench: 31.0% (31.0: 31.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vending-Bench 2: $5,396 (76.6: 76.6% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,284 (91.7: 46% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,479 (92.0: 46% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.2 (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-2 Index: 67.7 (leave-one-out 63.0 to 72.0) · Grade: C+ · Rank: #19 (17 to 25) · Verification: third-party published - GPQA Diamond: 91.4% (93.8: 93.8% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FEjo6Fy7XYxUDcCwbejfcYX.eval - Humanity's Last Exam: 27.8% (59.8: 59.8% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 52.9% (55.7: 55.7% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 86.2% (87.5: 87.5% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 45.8% (55.9: 55.9% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - SimpleQA Verified: 37.1% (49.1: 49.1% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FE3325ytkSNYP8ez53VA4fE.eval - SWE-bench Verified: 73.8% (88.4: 88.4% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F9owdbpnWJ5pkQQC3GWKP5L.eval - WeirdML: 72.2% (77.7: 77.7% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,416 (20.1: 10% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 96.1% (96.1: 96.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FLiUMZyNSM8q93KDubWgYCh.eval - FrontierMath (Tiers 1–3): 67.4% (71.9: 71.9% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 31.7% (32.5: 32.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - ProofBench: 15.0% (15.0: 15.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 64.9% (76.6: 76.6% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - GDPval: 49.7% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gdpval - METR Time Horizon: 5.9 h (84.4: 84.4% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/research - APEX-Agents: 34.4% (72.6: 72.6% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $3,591 (63.4: 63.4% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 41.1% (74.3: 74.3% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,244 (80.4: 40% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,438 (80.3: 40% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### DeepSeek V4 Pro (DeepSeek, China) URL: https://opencharts.com/benchmarks/deepseek-v4-pro Index: 67.5 (leave-one-out 63.3 to 71.2) · Grade: C+ · Rank: #20 (17 to 23) · Verification: third-party published - GPQA Diamond: 91.7% (94.2: 94.2% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 61.3% (64.5: 64.5% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 90.5% (91.9: 91.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 18.0% (55.7: 55.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 52.9% (70.0: 70.0% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F2xfFf6KZAp2gUWYhHGkFgs.eval - SWE-bench Verified: 77.6% (93.0: 93.0% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FWUk3f44JhxtAJaSxi6gFZR.eval - SciCode: 50.0% (80.6: 80.6% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 66.2% (71.3: 71.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 17.6% (33.0: 33.0% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Code Arena (WebDev): 1,582 (44.9: 22% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 98.6% (98.6: 98.6% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 64.6% (68.9: 68.9% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fb6yWmeFgsWeEiG9Qzs3NSf.eval - FrontierMath Tier 4: 26.8% (27.5: 27.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fa4uV8HMMBAcHXLVSmZfBgX.eval - ProofBench: 50.0% (50.0: 50.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vending-Bench 2: $3,285 (60.6: 60.6% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,460 (86.4: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Grok 4.5 (xAI, United States) URL: https://opencharts.com/benchmarks/grok-4-5 Index: 64.2 (leave-one-out 59.3 to 67.4) · Grade: C · Rank: #21 (21 to 28) · Verification: third-party published - GPQA Diamond: 93.4% (96.7: 96.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 52.6% (55.4: 55.4% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 87.2% (88.5: 88.5% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 70.0% (85.5: 85.5% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 15.4% (47.8: 47.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 48.3% (63.9: 63.9% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F7HWW3sjRstKzb9JpcnvLaV.eval - SciCode: 54.1% (87.1: 87.1% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 46.4% (50.0: 50.0% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 42.4% (79.4: 79.4% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Code Arena (WebDev): 1,556 (39.9: 20% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 97.8% (97.8: 97.8% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 57.2% (61.0: 61.0% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 24.4% (25.0: 25.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - ProofBench: 31.0% (31.0: 31.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 34.2% (72.2: 72.2% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $3,887 (66.0: 66.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,282 (91.1: 46% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,471 (89.7: 45% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Sonnet 4.6 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-sonnet-4-6 Index: 63.9 (leave-one-out 61.0 to 69.7) · Grade: C · Rank: #22 (18 to 28) · Verification: third-party published - GPQA Diamond: 87.4% (88.1: 88.1% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F58eQmyCfPa3FufhXJFvAL2.eval - ARC-AGI-2: 60.4% (63.6: 63.6% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 86.5% (87.8: 87.8% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 3.1% (9.7: 9.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 35.5% (47.0: 47.0% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SWE-bench Verified: 75.2% (90.1: 90.1% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FKEdgiMmtJKHxgGiS82agYj.eval - SciCode: 46.8% (75.4: 75.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 66.1% (71.1: 71.1% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 24.3% (45.5: 45.5% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 49.0% (66.8: 66.8% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,521 (34.0: 17% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 85.8% (85.8: 85.8% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F9LWdzVN5w4ihNvncsaCWMu.eval - ProofBench: 45.0% (45.0: 45.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 53.4% (63.0: 63.0% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - OSWorld 2.0: 9.3% (29.6: 29.6% of the best (31.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://osworld-v2.xlang.ai/ - APEX-Agents: 23.7% (50.0: 50.0% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $7,204 (85.9: 85.9% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 54.9% (99.3: 99.3% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,275 (89.2: 45% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,472 (90.0: 45% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### DeepSeek V4 Flash (DeepSeek, China) URL: https://opencharts.com/benchmarks/deepseek-v4-flash Index: 63.7 (leave-one-out 59.8 to 67.5) · Grade: C · Rank: #23 (20 to 28) · Verification: third-party published - GPQA Diamond: 91.0% (93.3: 93.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 61.4% (64.6: 64.6% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 89.0% (90.4: 90.4% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 61.1% (74.6: 74.6% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 16.6% (51.3: 51.3% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 33.6% (44.5: 44.5% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FTqFmiGvqozt7TSZYmhgF8m.eval - SciCode: 49.9% (80.4: 80.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 63.0% (67.8: 67.8% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 18.8% (35.2: 35.2% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Code Arena (WebDev): 1,580 (44.5: 22% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 94.4% (94.4: 94.4% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 57.5% (61.4: 61.4% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmKhVMEicZeGjdvJdTRd7mK.eval - FrontierMath Tier 4: 24.4% (25.0: 25.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FWYs83xB6NwrkMNyQJtuMcQ.eval - ProofBench: 56.0% (56.0: 56.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Text Arena: 1,438 (80.5: 40% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GLM 5.2 (Zhipu AI, China) URL: https://opencharts.com/benchmarks/glm-5-2 Index: 63.3 (leave-one-out 59.6 to 66.9) · Grade: C · Rank: #24 (21 to 28) · Verification: third-party published - GPQA Diamond: 91.9% (94.5: 94.5% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 22.8% (24.0: 24.0% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 77.0% (78.2: 78.2% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 58.8% (71.8: 71.8% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 20.9% (64.6: 64.6% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 34.2% (45.2: 45.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FWgjFZzojpiv54XqUhvfuUG.eval - SWE-bench Verified: 78.7% (94.3: 94.3% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F4AfhcmYVNrw6gM5u8CQsyy.eval - SciCode: 50.5% (81.3: 81.3% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 70.1% (75.5: 75.5% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 24.5% (45.8: 45.8% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 55.0% (74.9: 74.9% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,587 (45.9: 23% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 86.4% (86.4: 86.4% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 59.2% (63.2: 63.2% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 29.3% (30.0: 30.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FJebA8dYLVov6H49ZQSLXh5.eval - ProofBench: 35.0% (35.0: 35.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 35.6% (75.1: 75.1% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $8,314 (90.5: 90.5% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,472 (89.8: 45% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 3.6 Flash (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-6-flash Index: 63.2 (leave-one-out 58.7 to 67.5) · Grade: C · Rank: #25 (20 to 28) · Verification: third-party published - GPQA Diamond: 94.1% (97.7: 97.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 60.4% (63.6: 63.6% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 91.2% (92.6: 92.6% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 10.6% (32.7: 32.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 66.2% (87.6: 87.6% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FF83pa7Gfsb4UNVttTnS5Hk.eval - SciCode: 52.7% (84.9: 84.9% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 56.1% (60.4: 60.4% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 34.4% (64.3: 64.3% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 53.5% (72.9: 72.9% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,538 (36.8: 18% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 94.2% (94.2: 94.2% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 58.9% (62.9: 62.9% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fnb9LnDeeQyvRCqsbsTANGD.eval - FrontierMath Tier 4: 22.0% (22.5: 22.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FJCMjXU6FRiiLV3xvs9R2Xv.eval - ProofBench: 36.0% (36.0: 36.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,285 (91.9: 46% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,480 (92.3: 46% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 3 Flash (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-flash Index: 62.5 (leave-one-out 57.0 to 66.6) · Grade: C · Rank: #26 (21 to 30) · Verification: third-party published - GPQA Diamond: 89.4% (91.0: 91.0% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FgVdraDT3pXmaExP9VbqME4.eval - ARC-AGI-2: 33.6% (35.4: 35.4% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 84.7% (86.0: 86.0% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 61.1% (74.6: 74.6% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - SimpleQA Verified: 66.8% (88.4: 88.4% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F7cc6cDCXjsPjtKt36bbeMR.eval - SWE-bench Verified: 75.4% (90.3: 90.3% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjxoyMNmTnYirjLJN6wRNmG.eval - WeirdML: 61.6% (66.3: 66.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,438 (22.4: 11% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 95.6% (95.6: 95.6% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fa4uE6KgMhGSUjcnemRnbkA.eval - FrontierMath (Tiers 1–3): 51.2% (54.7: 54.7% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FEtgrDq6ei6kZe8K7QbXUe3.eval - FrontierMath Tier 4: 17.1% (17.5: 17.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FZczEhu69FQ9Cnns8NT6ew4.eval - ProofBench: 15.0% (15.0: 15.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 64.3% (75.9: 75.9% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - APEX-Agents: 24.0% (50.6: 50.6% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $3,635 (63.8: 63.8% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 49.8% (90.1: 90.1% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,272 (88.4: 44% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,474 (90.4: 45% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Opus 4.5 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-opus-4-5 Index: 62.4 (leave-one-out 57.4 to 66.4) · Grade: C · Rank: #27 (21 to 30) · Verification: third-party published - GPQA Diamond: 86.0% (86.3: 86.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FKwwo3UFcPDk6iLAeSeeRaa.eval - Humanity's Last Exam: 25.2% (54.2: 54.2% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 37.6% (39.6: 39.6% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 80.0% (81.2: 81.2% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 62.0% (75.7: 75.7% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - SimpleQA Verified: 45.7% (60.4: 60.4% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FCA5LW4phRnHSsZUGXEm8H6.eval - SWE-bench Verified: 76.7% (91.8: 91.8% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmKQKxkpGDrFHBPBYR3mBrB.eval - WeirdML: 63.7% (68.6: 68.6% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,495 (29.8: 15% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 86.1% (86.1: 86.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FikUWdu6HJ2qRgZynfzDCFe.eval - FrontierMath (Tiers 1–3): 34.4% (36.7: 36.7% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmS5EwCRZVpvUvJEehPUDZu.eval - FrontierMath Tier 4: 4.9% (5.0: 5.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fazk4theXxAsizmoi2HDD49.eval - ProofBench: 36.0% (36.0: 36.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 63.1% (74.5: 74.5% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - GDPval: 45.5% (91.5: 91.5% of the best (49.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gdpval - METR Time Horizon: 4.9 h (81.7: 81.7% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/research - APEX-Agents: 20.7% (43.7: 43.7% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $4,967 (73.9: 73.9% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Cybench: 82.0% (88.2: 88.2% of the best (93.0%)) — developer-reported; source: Epoch AI — AI Benchmarking Hub, https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf - DeepResearch Bench: 54.8% (99.1: 99.1% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Text Arena: 1,473 (90.2: 45% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5 (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5 Index: 62.0 (leave-one-out 59.4 to 65.0) · Grade: C · Rank: #28 (21 to 28) · Verification: third-party published - GPQA Diamond: 86.2% (86.4: 86.4% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F9nPyNrZxwtoT7eS4DZbQF6.eval - Humanity's Last Exam: 25.3% (54.5: 54.5% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 9.9% (10.4: 10.4% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 65.7% (66.7: 66.7% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 56.7% (69.2: 69.2% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - CritPt: 12.6% (39.0: 39.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 50.1% (66.3: 66.3% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FgTLurbyKU63i2opsGkJGwh.eval - SWE-bench Verified: 73.6% (88.1: 88.1% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FC85Bnci9xVzq5jmt8zNnbw.eval - Aider Polyglot: 88.0% (100.0: best published result) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://aider.chat/docs/leaderboards/#polyglot-leaderboard - SciCode: 42.9% (69.2: 69.2% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 60.7% (65.3: 65.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,419 (20.4: 10% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 91.4% (91.4: 91.4% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F3D8aaFBLZwQcnWHYKwvnmq.eval - MATH Level 5: 98.1% (100.0: best published result) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/math-level-5 - FrontierMath (Tiers 1–3): 55.4% (59.2: 59.2% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 22.0% (22.5: 22.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - ProofBench: 18.0% (18.0: 18.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 49.6% (58.6: 58.6% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - GDPval: 34.8% (70.0: 70.0% of the best (49.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gdpval - METR Time Horizon: 3.4 h (76.4: 76.4% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/research - APEX-Agents: 18.3% (38.6: 38.6% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - DeepResearch Bench: 49.6% (89.7: 89.7% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://drb.futuresearch.ai - Vision Arena: 1,210 (71.3: 36% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,434 (79.4: 40% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text - Fiction.LiveBench: 96.9% (96.9: 96.9% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://fiction.live/stories/Fiction-liveBench-Mar-14-2025/oQdzQvKHw8JyXbN87/home ### Kimi K2.7 Code (Moonshot AI, China) URL: https://opencharts.com/benchmarks/kimi-k2-7-code Index: 58.6 (leave-one-out 56.2 to 61.2) · Grade: C- · Rank: #29 (29 to 30) · Verification: third-party published - GPQA Diamond: 87.9% (88.8: 88.8% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FXneEig6pACkAWeekvgCwNk.eval - SimpleBench: 57.9% (70.7: 70.7% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 10.0% (31.0: 31.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 36.5% (48.3: 48.3% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F5AijYw5bFJtSuahnbzNo6Z.eval - SciCode: 47.5% (76.5: 76.5% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 54.1% (58.3: 58.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 30.1% (56.2: 56.2% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - CursorBench: 49.7% (67.7: 67.7% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,472 (26.7: 13% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 95.6% (95.6: 95.6% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FENc5ZM5TpBYNKkTWN4xyQM.eval - FrontierMath (Tiers 1–3): 54.0% (57.7: 57.7% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FX3RxRkt2pBUXYJjAQHsUo6.eval - FrontierMath Tier 4: 12.2% (12.5: 12.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FUTVSgirxQpTgmTHbgE9EUk.eval - APEX-Agents: 27.6% (58.2: 58.2% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $5,083 (74.6: 74.6% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 ### GLM 5.1 (Zhipu AI, China) URL: https://opencharts.com/benchmarks/glm-5-1 Index: 57.5 (leave-one-out 54.7 to 60.5) · Grade: C- · Rank: #30 (29 to 30) · Verification: third-party published - GPQA Diamond: 89.9% (91.7: 91.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FCBbeMKRRSeRsXnvSfn7Png.eval - SimpleBench: 55.1% (67.3: 67.3% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 4.6% (14.2: 14.2% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 34.0% (45.0: 45.0% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FR2fMFAtRXah864hMifdQGV.eval - SWE-bench Verified: 74.2% (88.9: 88.9% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/swe-bench-verified - SciCode: 43.8% (70.5: 70.5% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 57.1% (61.5: 61.5% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,508 (31.9: 16% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 93.3% (93.3: 93.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fb7Lw9ExxAhVY3hdExoh4MP.eval - FrontierMath (Tiers 1–3): 36.8% (39.3: 39.3% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FMhwme9odWJMzMjoYkW8r6B.eval - ProofBench: 22.2% (22.2: 22.2% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vending-Bench 2: $5,634 (77.9: 77.9% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,466 (88.2: 44% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Sonnet 4.5 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-sonnet-4-5 Index: 54.3 (leave-one-out 50.9 to 57.0) · Grade: C- · Rank: #31 (31 to 32) · Verification: third-party published - GPQA Diamond: 82.3% (81.0: 81.0% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FXaGuWDxPtqVvyxqdMncKFj.eval - Humanity's Last Exam: 13.7% (29.5: 29.5% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 13.6% (14.3: 14.3% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 63.7% (64.6: 64.6% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 54.3% (66.3: 66.3% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - CritPt: 1.1% (3.5: 3.5% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 30.7% (40.6: 40.6% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FXQEGUQjpXupjdyKoGuSmum.eval - SWE-bench Verified: 71.3% (85.4: 85.4% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FLti64yMEsBWSzgEMFCSVhp.eval - SciCode: 44.7% (72.0: 72.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 47.7% (51.4: 51.4% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,392 (17.7: 9% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 77.8% (77.8: 77.8% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fk5PDL6MpCC6P5eLAhodJiS.eval - MATH Level 5: 97.7% (99.6: 99.6% of the best (98.1%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/math-level-5 - FrontierMath (Tiers 1–3): 23.9% (25.5: 25.5% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F7oUkeLfwMK3MtkSNHX6QPu.eval - FrontierMath Tier 4: 2.4% (2.5: 2.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fd3m9Xip8cof4s77z6LKkEh.eval - ProofBench: 19.0% (19.0: 19.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 46.5% (54.9: 54.9% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - GDPval: 42.5% (85.5: 85.5% of the best (49.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gdpval - METR Time Horizon: 2.0 h (69.1: 69.1% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ - Vending-Bench 2: $3,839 (65.6: 65.6% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Cybench: 60.0% (64.5: 64.5% of the best (93.0%)) — developer-reported; source: Epoch AI — AI Benchmarking Hub, https://assets.anthropic.com/m/64823ba7485345a7/Claude-Opus-4-5-System-Card.pdf - DeepResearch Bench: 52.6% (95.1: 95.1% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://drb.futuresearch.ai - Text Arena: 1,456 (85.5: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.1 (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-1 Index: 54.0 (leave-one-out 51.5 to 56.4) · Grade: C- · Rank: #32 (31 to 32) · Verification: third-party published - GPQA Diamond: 87.6% (88.5: 88.5% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FCSVtMksEVjBNsGJGjW4GcQ.eval - Humanity's Last Exam: 23.7% (50.9: 50.9% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 17.6% (18.6: 18.6% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 72.8% (73.9: 73.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - LiveBench: 78.8% (95.7: 95.7% of the best (82.3%)) [not counted: too few models on this test] — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://livebench.ai/#/ - SimpleBench: 53.2% (65.0: 65.0% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - CritPt: 4.9% (15.0: 15.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 48.0% (63.5: 63.5% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FQCLnRy69NBH9ip3VYtQneh.eval - SWE-bench Verified: 68.0% (81.4: 81.4% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FYoHtZpMNdsRE6qdj2K4Ucj.eval - SciCode: 43.3% (69.8: 69.8% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 60.8% (65.4: 65.4% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,392 (17.7: 9% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 88.6% (88.6: 88.6% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmYKKTdV27TRJv9FpUg7rkn.eval - Terminal-Bench: 47.6% (56.2: 56.2% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - APEX-Agents: 17.5% (36.9: 36.9% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $1,473 (34.8: 34.8% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 42.8% (77.4: 77.4% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,250 (82.0: 41% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,455 (85.1: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Opus 4.1 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-opus-4-1 Index: 50.7 (leave-one-out 47.4 to 62.0) · Grade: C- · Rank: #33 (28 to 34) · Verification: third-party published - GPQA Diamond: 77.3% (73.9: 73.9% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FerLgyKQs2mg2memE9SJvK5.eval - Humanity's Last Exam: 11.5% (24.8: 24.8% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - SimpleBench: 60.0% (73.3: 73.3% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - SWE-bench Verified: 73.3% (87.9: 87.9% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FcA6Mq67yaynLhY8F5kes3c.eval - WeirdML: 45.9% (49.4: 49.4% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,389 (17.4: 9% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 68.9% (68.9: 68.9% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FUH9RY8MCAotDgFocHKA4j6.eval - FrontierMath (Tiers 1–3): 12.6% (13.5: 13.5% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F5eRkmXCURH2pRbhJoukfnC.eval - FrontierMath Tier 4: 2.4% (2.5: 2.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F9zT72xRR6wZhiwraJTNPyz.eval - Terminal-Bench: 38.0% (44.9: 44.9% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - GDPval: 43.6% (87.7: 87.7% of the best (49.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gdpval - METR Time Horizon: 114 min (68.1: 68.1% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ - Cybench: 42.0% (45.2: 45.2% of the best (93.0%)) — developer-reported; source: Epoch AI — AI Benchmarking Hub, https://cybench.github.io/ - DeepResearch Bench: 48.3% (87.3: 87.3% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://drb.futuresearch.ai - Text Arena: 1,450 (83.6: 42% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.4 mini (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-4-mini Index: 49.3 (leave-one-out 44.3 to 53.0) · Grade: D · Rank: #34 (33 to 34) · Verification: third-party published - GPQA Diamond: 86.9% (87.4: 87.4% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FgKB2rCD7mHzFN2MCt2YCni.eval - ARC-AGI-2: 18.9% (19.9: 19.9% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 63.7% (64.6: 64.6% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 10.0% (31.0: 31.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 29.4% (38.9: 38.9% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FZ3pjKgAULY9j9wmzQpdqZG.eval - SciCode: 49.9% (80.4: 80.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 60.3% (64.9: 64.9% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 27.0% (50.6: 50.6% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Code Arena (WebDev): 1,397 (18.2: 9% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 88.9% (88.9: 88.9% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FQGMk6JNCYuUkFMF4zdVCsM.eval - FrontierMath (Tiers 1–3): 51.2% (54.7: 54.7% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 9.8% (10.0: 10.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - ProofBench: 21.0% (21.0: 21.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 24.6% (51.9: 51.9% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - DeepResearch Bench: 36.3% (65.6: 65.6% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,252 (82.6: 41% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,448 (83.2: 42% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Inkling (Thinking Machines, United States) URL: https://opencharts.com/benchmarks/inkling Index: 42.8 (leave-one-out 36.5 to 46.4) · Grade: D · Rank: #35 (35 to 36) · Verification: third-party published - GPQA Diamond: 88.3% (89.4: 89.4% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 36.5% (38.5: 38.5% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 79.5% (80.7: 80.7% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 50.0% (61.1: 61.1% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 5.4% (16.8: 16.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 40.3% (53.3: 53.3% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FbpnzC6CCi3hWuKCfqX4KEY.eval - SciCode: 46.1% (74.3: 74.3% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 32.3% (34.8: 34.8% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 14.0% (26.2: 26.2% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Code Arena (WebDev): 1,409 (19.3: 10% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 88.9% (88.9: 88.9% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 33.3% (35.6: 35.6% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F6KvQvuMMX5zAuteUc6pFcW.eval - FrontierMath Tier 4: 4.9% (5.0: 5.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjvJPoRYkPH9Sb5MX252oHp.eval - ProofBench: 0.0% (0.0: 0.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Text Arena: 1,439 (80.7: 40% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 2.5 Pro (Google, United States) URL: https://opencharts.com/benchmarks/gemini-2-5-pro Index: 39.5 (leave-one-out 35.2 to 41.9) · Grade: F · Rank: #36 · Verification: third-party published - GPQA Diamond: 85.3% (85.2: 85.2% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FcqUZ6gcq2HBnShFRWvKjGr.eval - ARC-AGI-2: 4.9% (5.1: 5.1% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 41.0% (41.6: 41.6% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 2.0% (6.2: 6.2% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SWE-bench Verified: 57.6% (69.0: 69.0% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FKL724uHFGbiXj4o5PHyLn2.eval - SciCode: 42.8% (69.0: 69.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 54.0% (58.2: 58.2% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,226 (7.2: 4% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 84.2% (84.2: 84.2% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FiJp7PgYZa3PiXAV8UWd7MC.eval - FrontierMath (Tiers 1–3): 24.6% (26.2: 26.2% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FFkywQ2KS4AkLPzjSPmhHcS.eval - FrontierMath Tier 4: 0.0% (0.0: 0.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F2z7kL9DpfvY44T9Wo99FYx.eval - Terminal-Bench: 32.6% (38.5: 38.5% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - GDPval: 23.3% (46.9: 46.9% of the best (49.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gdpval - APEX-Agents: 6.6% (13.9: 13.9% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $574 (4.4: 4.4% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 41.5% (75.0: 75.0% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,246 (81.1: 41% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,446 (82.5: 41% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.5 Pro (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-5-pro Index: 93.9 (leave-one-out 92.5 to 95.5) · Grade: A · Rank: provisional · Verification: third-party published - ARC-AGI-2: 84.6% (89.0: 89.0% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 96.5% (98.0: 98.0% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 76.9% (93.9: 93.9% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lmcouncil.ai/benchmarks - CritPt: 30.6% (94.6: 94.6% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - FrontierMath (Tiers 1–3): 87.7% (93.6: 93.6% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 78.0% (80.0: 80.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath ### GPT-5.4 Pro (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-4-pro Index: 93.4 (leave-one-out 92.5 to 94.6) · Grade: A · Rank: provisional · Verification: third-party published - GPQA Diamond: 94.6% (98.3: 98.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - Humanity's Last Exam: 44.3% (95.3: 95.3% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 83.3% (87.7: 87.7% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 94.5% (95.9: 95.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 74.1% (90.5: 90.5% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - CritPt: 30.0% (92.9: 92.9% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 46.3% (61.2: 61.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FRfSJbuuuMkeVHQj7xrth4u.eval - WeirdML: 57.4% (61.8: 61.8% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierMath (Tiers 1–3): 82.5% (88.0: 88.0% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FZFVa8dw7oiir6PhcF8XZMo.eval - FrontierMath Tier 4: 58.5% (60.0: 60.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FW6b3C8AeNC7CAFs5DQma8N.eval ### ERNIE 5.1 (Baidu, China) URL: https://opencharts.com/benchmarks/ernie-5-1 Index: 88.8 · Grade: A- · Rank: provisional · Verification: third-party published - Text Arena: 1,468 (88.8: 44% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 3 Deep Think (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-deep-think Index: 88.7 (leave-one-out 84.3 to 93.3) · Grade: A- · Rank: provisional · Verification: third-party published - ARC-AGI-2: 84.6% (89.0: 89.0% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 96.0% (97.5: 97.5% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 25.7% (79.6: 79.6% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt ### GPT-5.6 Sol Pro (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-6-sol-pro Index: 87.7 (leave-one-out 84.8 to 89.4) · Grade: A- · Rank: provisional · Verification: third-party published - SimpleBench: 71.7% (87.5: 87.5% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - WeirdML: 89.4% (96.3: 96.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierMath Tier 4: 80.5% (82.5: 82.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FcdWwJSH3jVfrKyhUHmT54T.eval - APEX-Agents: 40.0% (84.4: 84.4% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents ### Seed 2.0 Pro (ByteDance, China) URL: https://opencharts.com/benchmarks/seed-2-0-pro Index: 84.7 (leave-one-out 84.0 to 85.3) · Grade: B+ · Rank: provisional · Verification: third-party published - Vision Arena: 1,257 (84.0: 42% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,456 (85.3: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.3 Codex (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-3-codex Index: 80.9 (leave-one-out 70.9 to 80.9) · Grade: B+ · Rank: provisional · Verification: third-party published - SWE-bench Verified: 74.8% (89.6: 89.6% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F6w5buMbgeEqtRoE55cLZVv.eval - WeirdML: 79.3% (85.4: 85.4% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,409 (19.3: 10% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Terminal-Bench: 78.4% (92.6: 92.6% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - METR Time Horizon: 5.8 h (84.2: 84.2% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/research - APEX-Agents: 31.8% (67.1: 67.1% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $5,940 (79.6: 79.6% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 ### Qwen3.6 Max (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-6-max Index: 76.3 (leave-one-out 70.9 to 81.6) · Grade: B · Rank: provisional · Verification: third-party published - GPQA Diamond: 87.4% (88.1: 88.1% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FJv2hJvjpvA8xJW3Bzv7bfw.eval - SimpleBench: 63.0% (76.9: 76.9% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - SimpleQA Verified: 52.0% (68.8: 68.8% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FiyPiR8GUuyNNA6Fssziq32.eval - SWE-bench Verified: 76.7% (91.8: 91.8% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FFkDMVLdEnJDR5PyUCT6BNQ.eval - Code Arena (WebDev): 1,479 (27.6: 14% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 91.1% (91.1: 91.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FEfJpx5HztMmL5EQSwsyqMe.eval - Vending-Bench 2: $4,254 (68.9: 68.9% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,460 (86.5: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Muse Spark 1.2 (Meta, United States) URL: https://opencharts.com/benchmarks/muse-spark-1-2 Index: 75.2 (leave-one-out 70.7 to 81.6) · Grade: B · Rank: provisional · Verification: third-party published - SimpleBench: 74.5% (91.0: 91.0% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 17.7% (54.8: 54.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 60.3% (79.8: 79.8% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmLLmVVvo6t23MM5yxq7ryW.eval - SciCode: 56.4% (90.9: 90.9% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 60.3% (64.9: 64.9% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,534 (36.1: 18% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 43.0% (43.0: 43.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,292 (93.9: 47% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,499 (97.6: 49% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.2 Pro (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-2-pro Index: 73.0 (leave-one-out 63.4 to 73.0) · Grade: B- · Rank: provisional · Verification: third-party published - ARC-AGI-2: 54.2% (57.0: 57.0% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 90.5% (91.9: 91.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 57.4% (70.1: 70.1% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - FrontierMath (Tiers 1–3): 74.0% (79.0: 79.0% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 46.0% (47.1: 47.1% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath ### Gemini 3 Pro (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-pro Index: 72.3 (leave-one-out 69.4 to 76.9) · Grade: B- · Rank: provisional · Verification: third-party published - GPQA Diamond: 92.6% (95.5: 95.5% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FCg54udQGcTCnfDGX2qePLZ.eval - Humanity's Last Exam: 37.5% (80.7: 80.7% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 31.1% (32.7: 32.7% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 75.0% (76.1: 76.1% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 76.4% (93.3: 93.3% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - CritPt: 6.9% (21.4: 21.4% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SWE-bench Verified: 72.9% (87.4: 87.4% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F5XdYZBG5Wy8DUMRqcYv4fn.eval - WeirdML: 69.9% (75.3: 75.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,439 (22.5: 11% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 91.4% (91.4: 91.4% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F4QM4h7ndQiDRL2anjKeddK.eval - ProofBench: 20.0% (20.0: 20.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 69.4% (81.9: 81.9% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - GDPval: 40.3% (81.1: 81.1% of the best (49.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gdpval - METR Time Horizon: 3.7 h (77.9: 77.9% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/research - APEX-Agents: 31.5% (66.5: 66.5% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $5,478 (77.0: 77.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 46.3% (83.7: 83.7% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,289 (93.0: 47% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,486 (93.8: 47% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Muse Spark (Meta, United States) URL: https://opencharts.com/benchmarks/muse-spark Index: 71.3 (leave-one-out 71.3 to 82.9) · Grade: B- · Rank: provisional · Verification: third-party published - GPQA Diamond: 89.8% (91.6: 91.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - Humanity's Last Exam: 40.6% (87.2: 87.2% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - CritPt: 11.3% (35.1: 35.1% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 51.5% (83.0: 83.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Mock AIME 2024–2025: 88.9% (88.9: 88.9% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - ProofBench: 17.0% (17.0: 17.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,294 (94.4: 47% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,488 (94.5: 47% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Muse Spark 1.1 (Meta, United States) URL: https://opencharts.com/benchmarks/muse-spark-1-1 Index: 71.3 (leave-one-out 67.3 to 76.7) · Grade: B- · Rank: provisional · Verification: third-party published - CritPt: 15.1% (46.7: 46.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 57.8% (76.4: 76.4% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FbCdwpRXgZtvQLYC4wA9JXr.eval - SciCode: 58.2% (93.8: 93.8% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,541 (37.3: 19% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 39.0% (39.0: 39.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 41.9% (88.4: 88.4% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $6,520 (82.6: 82.6% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,279 (90.4: 45% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,492 (95.7: 48% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.8 Max (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-8-max Index: 71.2 (leave-one-out 61.8 to 79.1) · Grade: B- · Rank: provisional · Verification: third-party published - GPQA Diamond: 92.7% (95.6: 95.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - CritPt: 20.0% (61.9: 61.9% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 45.8% (60.6: 60.6% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fbfftd5EzUK9yAekafEofvK.eval - SciCode: 52.9% (85.3: 85.3% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,686 (69.1: 35% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 99.4% (99.4: 99.4% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 74.7% (79.8: 79.8% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FABT89ieSWGnGj4RYMduHiG.eval - FrontierMath Tier 4: 46.3% (47.5: 47.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FPBRdMZyz4L4CqPpXyLD2Lk.eval - ProofBench: 58.0% (58.0: 58.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,300 (96.2: 48% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,480 (92.1: 46% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 3.8 Flash (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-8-flash Index: 68.9 (leave-one-out 59.8 to 75.8) · Grade: C+ · Rank: provisional · Verification: third-party published - CritPt: 18.3% (56.6: 56.6% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 54.4% (87.7: 87.7% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - CursorBench: 69.2% (94.3: 94.3% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,567 (42.1: 21% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 48.0% (48.0: 48.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Text Arena: 1,494 (96.2: 48% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.7 Flash (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-7-flash Index: 67.4 (leave-one-out 50.8 to 83.9) · Grade: C+ · Rank: provisional · Verification: third-party published - GPQA Diamond: 82.3% (81.0: 81.0% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fi6i3gQf3ng2aAipvowPpws.eval - Mock AIME 2024–2025: 86.7% (86.7: 86.7% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FTjoS7sNuqVTc4LVyGRo4wa.eval - FrontierMath (Tiers 1–3): 19.3% (20.6: 20.6% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FdoNrgGQY5SytQV22WRqw5L.eval ### Qwen3.5 397B-A17B (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-5-397b-a17b Index: 65.7 (leave-one-out 60.2 to 77.6) · Grade: C+ · Rank: provisional · Verification: third-party published - GPQA Diamond: 86.4% (86.7: 86.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FbXrmQqzF9sXbuZfRrSc7NP.eval - Code Arena (WebDev): 1,399 (18.3: 9% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 88.9% (88.9: 88.9% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FJYuFdyZZSiMvmSZNzMTywb.eval - FrontierMath (Tiers 1–3): 31.2% (33.3: 33.3% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FA4yUA9PfzRr92JAqDDo9ro.eval - Vision Arena: 1,247 (81.3: 41% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,441 (81.2: 41% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.7 Max (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-7-max Index: 65.0 (leave-one-out 56.4 to 70.1) · Grade: C+ · Rank: provisional · Verification: third-party published - GPQA Diamond: 90.9% (93.1: 93.1% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FW6mTanyjvXhasFBzHukv4S.eval - SimpleBench: 70.4% (86.0: 86.0% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 13.4% (41.6: 41.6% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 55.8% (73.8: 73.8% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FDMjUrG6eZEqwcQ4SqUoViu.eval - SWE-bench Verified: 77.3% (92.6: 92.6% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F83hhDhcT9UUrZ2NiUbWSUU.eval - SciCode: 48.8% (78.7: 78.7% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,517 (33.3: 17% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 95.6% (95.6: 95.6% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FD3YDUZSfEUDsS4frLbEKDu.eval - FrontierMath (Tiers 1–3): 64.6% (68.9: 68.9% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FRG2biwnmGybSVTVxxMdRWa.eval - FrontierMath Tier 4: 34.1% (35.0: 35.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FFHZCKbtfvpKGXU99vZFnXN.eval - ProofBench: 26.0% (26.0: 26.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Text Arena: 1,474 (90.4: 45% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3 Max (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-max Index: 64.2 (leave-one-out 51.6 to 64.2) · Grade: C · Rank: provisional · Verification: third-party published - GPQA Diamond: 72.6% (67.3: 67.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F46RCssznDADYeg8MuSgvsj.eval - SimpleQA Verified: 48.7% (64.5: 64.5% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FfXG2GPt5F25tHRTPhAoZ4K.eval - Mock AIME 2024–2025: 73.3% (73.3: 73.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FULeuPtwp3ZoyuXQYuuQ82i.eval - MATH Level 5: 97.1% (99.0: 99.0% of the best (98.1%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/math-level-5 - FrontierMath (Tiers 1–3): 18.9% (20.2: 20.2% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FNrTH3EYHjp4hPdxgpCPGTX.eval - Vending-Bench 2: $72 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,435 (79.4: 40% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.6 Plus (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-6-plus Index: 64.1 (leave-one-out 57.4 to 70.9) · Grade: C · Rank: provisional · Verification: third-party published - GPQA Diamond: 88.4% (89.6: 89.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Ff6ZdMtRpdfskGbrBCzrKSn.eval - CritPt: 2.9% (8.8: 8.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 44.1% (58.4: 58.4% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F4LWNorP8Fy8V8doxirNFcH.eval - SWE-bench Verified: 57.9% (69.3: 69.3% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F8HV4HP9EM5KuLSwzifuLye.eval - SciCode: 40.7% (65.7: 65.7% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,460 (25.1: 13% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 93.3% (93.3: 93.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjYiGAFfUQKJuRStPzU3PX3.eval - FrontierMath (Tiers 1–3): 38.2% (40.8: 40.8% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjbyVWnEvvFVNNc8AEF6WcQ.eval - Vending-Bench 2: $5,115 (74.8: 74.8% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,444 (81.9: 41% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Grok 4.20 (xAI, United States) URL: https://opencharts.com/benchmarks/grok-4-20 Index: 63.2 (leave-one-out 42.9 to 68.0) · Grade: C · Rank: provisional · Verification: third-party published - GPQA Diamond: 89.3% (90.9: 90.9% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 65.1% (68.6: 68.6% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 89.5% (90.9: 90.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleQA Verified: 30.2% (39.9: 39.9% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FGVm7khGvtdnQKGfkgAHD8q.eval - WeirdML: 52.3% (56.3: 56.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,374 (16.1: 8% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 92.2% (92.2: 92.2% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 44.9% (47.9: 47.9% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FFLDt53waLTeshC7PK9uVy2.eval - FrontierMath Tier 4: 17.1% (17.5: 17.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FSX5FCWMk3HF83gGsP3WFA6.eval - ProofBench: 14.0% (14.0: 14.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 57.3% (67.7: 67.7% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - Vending-Bench 2: $4,663 (71.9: 71.9% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,256 (83.7: 42% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,475 (90.6: 45% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemma 4 26B A4B (Google, United States) URL: https://opencharts.com/benchmarks/gemma-4-26b-a4b Index: 63.1 (leave-one-out 56.3 to 70.0) · Grade: C · Rank: provisional · Verification: third-party published - GPQA Diamond: 73.2% (68.2: 68.2% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FUsx9KaiXvHZU4qrzgc4sLg.eval - CritPt: 0.0% (0.0: 0.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 40.0% (64.6: 64.6% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 35.2% (37.9: 37.9% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,361 (15.0: 8% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 82.2% (82.2: 82.2% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmstwSH7noJfvhGcd6Lz5kd.eval - Vision Arena: 1,242 (79.7: 40% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,438 (80.5: 40% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GLM 5.3 (Zhipu AI, China) URL: https://opencharts.com/benchmarks/glm-5-3 Index: 60.9 (leave-one-out 50.8 to 71.2) · Grade: C · Rank: provisional · Verification: third-party published - GPQA Diamond: 90.9% (93.1: 93.1% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - CritPt: 19.1% (59.3: 59.3% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 41.0% (54.2: 54.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FTY8R8FV7zCwByEipkqyg2P.eval - SciCode: 56.5% (91.0: 91.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 75.4% (81.2: 81.2% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,609 (50.6: 25% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 91.1% (91.1: 91.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 68.8% (73.4: 73.4% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 29.3% (30.0: 30.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - ProofBench: 49.0% (49.0: 49.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vending-Bench 2: $8,164 (89.9: 89.9% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,482 (92.8: 46% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Kimi K2 (Moonshot AI, China) URL: https://opencharts.com/benchmarks/kimi-k2 Index: 60.6 (leave-one-out 55.9 to 64.6) · Grade: C · Rank: provisional · Verification: third-party published - GPQA Diamond: 84.2% (83.7: 83.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmuGecFpAg9wHd6gexWJ99L.eval - Aider Polyglot: 59.1% (67.2: 67.2% of the best (88.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://aider.chat/docs/leaderboards/ - WeirdML: 42.8% (46.1: 46.1% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,322 (12.2: 6% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 83.1% (83.1: 83.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F8BM7dCHaUZgcYJwZWiBhiQ.eval - Terminal-Bench: 35.7% (42.1: 42.1% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - METR Time Horizon: 54 min (57.4: 57.4% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ - APEX-Agents: 4.1% (8.6: 8.6% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Text Arena: 1,430 (78.2: 39% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text - Fiction.LiveBench: 40.6% (40.6: 40.6% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://fiction.live/stories/Fiction-liveBench-Mar-14-2025/oQdzQvKHw8JyXbN87/home ### Hunyuan 3 (Tencent, China) URL: https://opencharts.com/benchmarks/hunyuan-3 Index: 58.8 (leave-one-out 32.4 to 85.2) · Grade: C- · Rank: provisional · Verification: third-party published - Code Arena (WebDev): 1,512 (32.4: 16% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Text Arena: 1,455 (85.2: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5 Pro (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-pro Index: 58.5 (leave-one-out 52.9 to 71.5) · Grade: C- · Rank: provisional · Verification: third-party published - Humanity's Last Exam: 31.6% (68.0: 68.0% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 18.3% (19.3: 19.3% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 70.2% (71.3: 71.3% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 61.6% (75.2: 75.2% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - WeirdML: 60.4% (65.0: 65.0% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierMath (Tiers 1–3): 55.8% (59.6: 59.6% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FTFDQUeRKbWiEodRQRrKdpy.eval - FrontierMath Tier 4: 19.5% (20.0: 20.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FHfDahF4ZXTN6bjfDFwcHTU.eval ### Kimi K2.6 (Moonshot AI, China) URL: https://opencharts.com/benchmarks/kimi-k2-6 Index: 58.5 (leave-one-out 50.7 to 64.1) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 90.8% (93.0: 93.0% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FLjgJoDosTLkiGVT59smA7q.eval - CritPt: 8.0% (24.8: 24.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 34.9% (46.2: 46.2% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SWE-bench Verified: 76.7% (91.8: 91.8% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FbUPtAK7spEwvmiC8pi6afc.eval - SciCode: 53.5% (86.2: 86.2% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 55.9% (60.1: 60.1% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - CursorBench: 47.6% (64.9: 64.9% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,509 (32.0: 16% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 96.1% (96.1: 96.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FBTeRdhgt6TFBj6QkGX2j6P.eval - FrontierMath (Tiers 1–3): 57.2% (61.0: 61.0% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FBxrNfyMffkVUq4n7U5eK3Q.eval - FrontierMath Tier 4: 25.6% (26.3: 26.3% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FeVKA2eavdRAbSN6AnzyBxA.eval - ProofBench: 16.0% (16.0: 16.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - OSWorld 2.0: 4.6% (14.6: 14.6% of the best (31.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://osworld-v2.xlang.ai/ - APEX-Agents: 18.9% (39.9: 39.9% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $6,205 (81.0: 81.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,263 (85.6: 43% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,461 (86.7: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.7 Plus (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-7-plus Index: 58.4 (leave-one-out 52.8 to 68.3) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 87.9% (88.8: 88.8% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FURoWjw9Htc6QPwABVcrybw.eval - CritPt: 9.1% (28.3: 28.3% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 45.5% (73.3: 73.3% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - FrontierCode: 10.2% (19.1: 19.1% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Mock AIME 2024–2025: 93.3% (93.3: 93.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FJvdw6GAXiS6T79bkpbk6Ra.eval - FrontierMath (Tiers 1–3): 34.4% (36.7: 36.7% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FTt8tBXHoCEYqGLePWXkM8u.eval - OSWorld 2.0: 2.8% (8.9: 8.9% of the best (31.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://osworld-v2.xlang.ai/ - Vision Arena: 1,266 (86.5: 43% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,455 (85.2: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemma 4 31B (Google, United States) URL: https://opencharts.com/benchmarks/gemma-4-31b Index: 57.0 (leave-one-out 51.3 to 65.6) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 75.8% (71.7: 71.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FP5L8qbXo42vTvbU3UvTNjk.eval - CritPt: 1.4% (4.4: 4.4% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 10.4% (13.8: 13.8% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FD5cbXacPnJ8akq2jAtifEQ.eval - SciCode: 43.4% (70.0: 70.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 52.3% (56.3: 56.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,363 (15.2: 8% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 73.3% (73.3: 73.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjFJucVunfLgoQqoFvPVQWP.eval - Vision Arena: 1,261 (85.2: 43% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,451 (84.1: 42% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Grok 4 (xAI, United States) URL: https://opencharts.com/benchmarks/grok-4 Index: 56.3 (leave-one-out 51.9 to 63.7) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 87.0% (87.6: 87.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 16.0% (16.8: 16.8% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 66.7% (67.7: 67.7% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 60.5% (73.9: 73.9% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - Aider Polyglot: 79.6% (90.5: 90.5% of the best (88.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://aider.chat/docs/leaderboards/#polyglot-leaderboard - WeirdML: 45.7% (49.2: 49.2% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Mock AIME 2024–2025: 84.0% (84.0: 84.0% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - Terminal-Bench: 27.2% (32.1: 32.1% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - GDPval: 21.1% (42.5: 42.5% of the best (49.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gdpval - METR Time Horizon: 110 min (67.6: 67.6% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ - APEX-Agents: 15.2% (32.1: 32.1% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Cybench: 43.0% (46.2: 46.2% of the best (93.0%)) — developer-reported; source: Epoch AI — AI Benchmarking Hub, https://data.x.ai/2025-08-20-grok-4-model-card.pdf - DeepResearch Bench: 47.3% (85.5: 85.5% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://drb.futuresearch.ai - Vision Arena: 1,184 (64.4: 32% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,411 (73.0: 37% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text - Fiction.LiveBench: 96.9% (96.9: 96.9% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://fiction.live/stories/Fiction-liveBench-Mar-14-2025/oQdzQvKHw8JyXbN87/home ### Qwen3.6 27B (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-6-27b Index: 56.3 (leave-one-out 42.4 to 70.1) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 85.9% (86.0: 86.0% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FbuLf8Ne29xmzMZBRQkvrQG.eval - CritPt: 0.9% (2.7: 2.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 37.3% (60.1: 60.1% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Mock AIME 2024–2025: 91.1% (91.1: 91.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F452pB5w5ea6uy92e2q2NMv.eval - FrontierMath (Tiers 1–3): 35.1% (37.5: 37.5% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FVonoEhAsGSJ2FHU33qnhui.eval ### Qwen3.5 Plus (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-5-plus Index: 54.8 (leave-one-out 44.2 to 61.9) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 84.8% (84.6: 84.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmGXYSUNCMisvaRzhQp6g2L.eval - SimpleQA Verified: 25.4% (33.5: 33.5% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjrawnM2wHFbVfJFJsap4kS.eval - Mock AIME 2024–2025: 86.7% (86.7: 86.7% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FcFbntXuxUtfreULvt7vEt7.eval - APEX-Agents: 13.6% (28.7: 28.7% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $1 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 ### Qwen3.8 Flash Next (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-8-flash-next Index: 54.5 · Grade: C- · Rank: provisional · Verification: third-party published - Code Arena (WebDev): 1,626 (54.5: 27% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev ### Muse Spark 1.3 (Meta, United States) URL: https://opencharts.com/benchmarks/muse-spark-1-3 Index: 53.6 · Grade: C- · Rank: provisional · Verification: third-party published - Code Arena (WebDev): 1,622 (53.6: 27% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev ### MiniMax M3 (MiniMax, China) URL: https://opencharts.com/benchmarks/minimax-m3 Index: 53.5 (leave-one-out 52.1 to 58.9) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 90.9% (93.1: 93.1% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FhDESmyGFrbZDomjff3rqzD.eval - SimpleBench: 45.8% (55.9: 55.9% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 3.7% (11.5: 11.5% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 45.4% (73.1: 73.1% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - FrontierCode: 14.7% (27.5: 27.5% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Code Arena (WebDev): 1,487 (28.8: 14% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 71.1% (71.1: 71.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F79ZP6iwoBaF5t9a5VRqPnw.eval - ProofBench: 18.0% (18.0: 18.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - OSWorld 2.0: 4.6% (14.6: 14.6% of the best (31.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://osworld-v2.xlang.ai/ - Vending-Bench 2: $2,158 (47.1: 47.1% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,237 (78.5: 39% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,443 (81.7: 41% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5 Codex (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-codex Index: 53.0 (leave-one-out 47.3 to 55.5) · Grade: C- · Rank: provisional · Verification: third-party published - WeirdML: 54.5% (58.7: 58.7% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Terminal-Bench: 44.3% (52.3: 52.3% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - APEX-Agents: 20.1% (42.4: 42.4% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents ### Qwen3.5 35B-A3B (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-5-35b-a3b Index: 52.2 (leave-one-out 42.1 to 62.3) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 83.5% (82.6: 82.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FRWjuCuVGvaRd4vDjJJ6aP2.eval - CritPt: 0.6% (1.8: 1.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 29.3% (47.2: 47.2% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,250 (8.2: 4% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 70.0% (70.0: 70.0% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fh9aoDxyWEhybvARHpyuoco.eval - Text Arena: 1,395 (68.8: 34% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Kimi K2.5 (Moonshot AI, China) URL: https://opencharts.com/benchmarks/kimi-k2-5 Index: 52.1 (leave-one-out 48.0 to 56.4) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 87.6% (88.5: 88.5% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FFmqGUq87N4JcuSbo6Uu3U5.eval - Humanity's Last Exam: 24.4% (52.4: 52.4% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 11.8% (12.4: 12.4% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 65.3% (66.3: 66.3% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 46.8% (57.1: 57.1% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 3.1% (9.7: 9.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 34.3% (45.4: 45.4% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FRbXXg4JEbTSBnzw3CMSHnD.eval - SWE-bench Verified: 73.8% (88.4: 88.4% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F5wVjUb3yn6bdjPSfPEMuiZ.eval - SciCode: 49.0% (78.9: 78.9% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 45.6% (49.1: 49.1% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - CursorBench: 31.9% (43.5: 43.5% of the best (73.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cursor.com/cursorbench - Code Arena (WebDev): 1,436 (22.3: 11% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 92.2% (92.2: 92.2% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FmbfZvrSNoXSJqCyD4VKqLe.eval - Terminal-Bench: 43.2% (51.0: 51.0% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - APEX-Agents: 14.4% (30.4: 30.4% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $1,198 (28.1: 28.1% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,250 (81.9: 41% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,451 (83.9: 42% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text - Fiction.LiveBench: 78.1% (78.1: 78.1% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://fiction.live/stories/Fiction-liveBench-Mar-14-2025/oQdzQvKHw8JyXbN87/home ### GLM 5 (Zhipu AI, China) URL: https://opencharts.com/benchmarks/glm-5 Index: 51.1 (leave-one-out 38.5 to 66.4) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 87.8% (88.8: 88.8% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fn4hyz75SVpwPEisKVVjvNg.eval - ARC-AGI-2: 4.9% (5.1: 5.1% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 44.7% (45.4: 45.4% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 53.2% (65.0: 65.0% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - SWE-bench Verified: 72.1% (86.4: 86.4% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FhKkJEu4ueLiNWsTtrJ78PS.eval - WeirdML: 48.2% (51.9: 51.9% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,436 (22.2: 11% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 80.0% (80.0: 80.0% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fccbx7KQY9bTD563vzBXwSg.eval - Terminal-Bench: 52.4% (61.9: 61.9% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - APEX-Agents: 17.2% (36.3: 36.3% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $4,432 (70.2: 70.2% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,458 (85.9: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.8 27B (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-8-27b Index: 51.0 (leave-one-out 43.1 to 59.7) · Grade: C- · Rank: provisional · Verification: third-party published - CritPt: 5.4% (16.8: 16.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 44.7% (72.0: 72.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,594 (47.4: 24% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 16.0% (16.0: 16.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,251 (82.5: 41% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,436 (79.8: 40% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Inkling Small (Thinking Machines, United States) URL: https://opencharts.com/benchmarks/inkling-small Index: 50.8 (leave-one-out 42.6 to 56.6) · Grade: C- · Rank: provisional · Verification: third-party published - GPQA Diamond: 88.5% (89.7: 89.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - ARC-AGI-2: 40.1% (42.3: 42.3% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 84.0% (85.3: 85.3% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 8.3% (25.7: 25.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 19.1% (25.3: 25.3% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FeJfbDKqeg9kNXFoxpeV7HG.eval - SciCode: 48.7% (78.5: 78.5% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,405 (19.0: 10% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 90.0% (90.0: 90.0% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 46.3% (49.4: 49.4% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FB6bZbqickq8o2fQcgDmzMo.eval - FrontierMath Tier 4: 17.1% (17.5: 17.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FDVD8G8DHNyzkowBzK7JvMt.eval - ProofBench: 6.0% (6.0: 6.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,207 (70.3: 35% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,407 (71.9: 36% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Nemotron 3 Ultra (NVIDIA, United States) URL: https://opencharts.com/benchmarks/nemotron-3-ultra Index: 49.8 (leave-one-out 41.3 to 58.2) · Grade: D · Rank: provisional · Verification: third-party published - GPQA Diamond: 85.4% (85.3: 85.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FbpZy8tAbh7UV6A6U7XZtBg.eval - CritPt: 3.1% (9.7: 9.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 39.9% (64.4: 64.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 43.5% (46.8: 46.8% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Mock AIME 2024–2025: 86.7% (86.7: 86.7% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FWYTWr3Xk4qxymhMsm7vRQ2.eval - ProofBench: 2.0% (2.0: 2.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 11.5% (24.3: 24.3% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Text Arena: 1,426 (77.1: 39% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GLM 4.7 (Zhipu AI, China) URL: https://opencharts.com/benchmarks/glm-4-7 Index: 48.6 (leave-one-out 47.3 to 53.8) · Grade: D · Rank: provisional · Verification: third-party published - GPQA Diamond: 83.3% (82.4: 82.4% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F6phXzd4aCbnBUSQmBsaywa.eval - SimpleBench: 47.7% (58.2: 58.2% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 1.7% (5.3: 5.3% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 32.2% (42.6: 42.6% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FaYau6EG7CCaWVSyVq3vvQc.eval - SciCode: 45.1% (72.8: 72.8% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,434 (22.1: 11% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 83.3% (83.3: 83.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FGYzmrzMfeSjNjm2J6L9PeT.eval - ProofBench: 6.0% (6.0: 6.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 33.4% (39.4: 39.4% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - APEX-Agents: 8.7% (18.4: 18.4% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $2,377 (50.2: 50.2% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,442 (81.4: 41% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### DeepSeek V3.1 Terminus (DeepSeek, China) URL: https://opencharts.com/benchmarks/deepseek-v3-1-terminus Index: 48.5 (leave-one-out 35.4 to 70.2) · Grade: D · Rank: provisional · Verification: third-party published - CritPt: 1.7% (5.3: 5.3% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 40.6% (65.5: 65.5% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Text Arena: 1,418 (74.8: 37% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### DeepSeek V3.2 (DeepSeek, China) URL: https://opencharts.com/benchmarks/deepseek-v3-2 Index: 48.3 (leave-one-out 43.5 to 53.3) · Grade: D · Rank: provisional · Verification: third-party published - GPQA Diamond: 83.4% (82.6: 82.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FCDdjbqiYYPtEBGoZLMPM3n.eval - ARC-AGI-2: 4.0% (4.2: 4.2% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 57.0% (57.9: 57.9% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleBench: 52.6% (64.2: 64.2% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - CritPt: 2.9% (8.8: 8.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - Aider Polyglot: 74.2% (84.3: 84.3% of the best (88.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://aider.chat/docs/leaderboards/#polyglot-leaderboard - SciCode: 38.9% (62.7: 62.7% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 46.7% (50.3: 50.3% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,360 (15.0: 8% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 87.8% (87.8: 87.8% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F8szQqvTsGG8gJ2TiFReRn9.eval - ProofBench: 8.0% (8.0: 8.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 39.6% (46.8: 46.8% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - APEX-Agents: 7.0% (14.8: 14.8% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $1,034 (23.4: 23.4% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,425 (76.8: 38% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GLM 5.3 Flash (Zhipu AI, China) URL: https://opencharts.com/benchmarks/glm-5-3-flash Index: 48.0 (leave-one-out 32.7 to 58.2) · Grade: D · Rank: provisional · Verification: third-party published - GPQA Diamond: 90.2% (92.1: 92.1% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - CritPt: 15.4% (47.8: 47.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 46.1% (74.3: 74.3% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,605 (49.7: 25% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 93.9% (93.9: 93.9% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 55.8% (59.6: 59.6% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - FrontierMath Tier 4: 17.1% (17.5: 17.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/frontiermath - ProofBench: 21.0% (21.0: 21.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,273 (88.4: 44% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,474 (90.5: 45% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.5 122B-A10B (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-5-122b-a10b Index: 47.3 (leave-one-out 37.9 to 62.2) · Grade: D · Rank: provisional · Verification: third-party published - CritPt: 0.9% (2.7: 2.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 35.6% (57.5: 57.5% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,358 (14.8: 7% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Vision Arena: 1,227 (75.6: 38% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,417 (74.7: 37% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Grok 4 Fast (xAI, United States) URL: https://opencharts.com/benchmarks/grok-4-fast Index: 47.1 (leave-one-out 40.1 to 51.4) · Grade: D · Rank: provisional · Verification: third-party published - ARC-AGI-2: 5.3% (5.6: 5.6% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 48.5% (49.2: 49.2% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - WeirdML: 42.9% (46.1: 46.1% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,162 (5.0: 3% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Cybench: 30.0% (32.3: 32.3% of the best (93.0%)) — developer-reported; source: Epoch AI — AI Benchmarking Hub, https://data.x.ai/2025-09-19-grok-4-fast-model-card.pdf - Text Arena: 1,418 (75.0: 38% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text - Fiction.LiveBench: 75.0% (75.0: 75.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://fiction.live/stories/Fiction-liveBench-Mar-14-2025/oQdzQvKHw8JyXbN87/home ### MiMo V2.5 (Xiaomi, China) URL: https://opencharts.com/benchmarks/mimo-v2-5 Index: 46.1 (leave-one-out 37.8 to 54.8) · Grade: D · Rank: provisional · Verification: third-party published - CritPt: 3.7% (11.5: 11.5% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 43.1% (69.4: 69.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,438 (22.4: 11% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 16.0% (16.0: 16.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,235 (77.9: 39% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,434 (79.2: 40% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.6 Flash (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-6-flash Index: 46.0 (leave-one-out 35.9 to 58.5) · Grade: D · Rank: provisional · Verification: third-party published - GPQA Diamond: 83.3% (82.4: 82.4% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FH2FvRXagQRQmiNa7WrNMU2.eval - SimpleBench: 35.2% (43.0: 43.0% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com - SimpleQA Verified: 15.9% (21.1: 21.1% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FFoMSS4FmH3VxpsvZiUJPCv.eval - Mock AIME 2024–2025: 84.4% (84.4: 84.4% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F9zx2EXPTvPmxsL9qFCjP3x.eval - FrontierMath (Tiers 1–3): 22.5% (24.0: 24.0% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FRoscWx4NUGVqcHdko6QhKd.eval ### Gemini 2.5 Flash (Google, United States) URL: https://opencharts.com/benchmarks/gemini-2-5-flash Index: 45.2 (leave-one-out 38.4 to 49.9) · Grade: D · Rank: provisional · Verification: third-party published - SimpleBench: 41.2% (50.3: 50.3% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - CritPt: 1.1% (3.4: 3.4% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - WeirdML: 41.9% (45.1: 45.1% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Terminal-Bench: 17.1% (20.2: 20.2% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - APEX-Agents: 1.8% (3.8: 3.8% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $549 (3.0: 3.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,214 (72.4: 36% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,410 (72.7: 36% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### MiMo V2.5 Pro (Xiaomi, China) URL: https://opencharts.com/benchmarks/mimo-v2-5-pro Index: 44.3 (leave-one-out 29.5 to 55.0) · Grade: D · Rank: provisional · Verification: third-party published - CritPt: 4.0% (12.4: 12.4% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 50.2% (81.0: 81.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,475 (27.1: 14% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 22.0% (22.0: 22.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Text Arena: 1,468 (88.8: 44% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Mistral Large 3 (Mistral AI, France) URL: https://opencharts.com/benchmarks/mistral-large-3 Index: 44.1 (leave-one-out 34.3 to 58.8) · Grade: D · Rank: provisional · Verification: third-party published - CritPt: 0.0% (0.0: 0.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 36.2% (58.4: 58.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,229 (7.3: 4% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Vision Arena: 1,205 (69.9: 35% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,414 (73.7: 37% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.5 27B (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-5-27b Index: 43.6 (leave-one-out 33.6 to 58.1) · Grade: D · Rank: provisional · Verification: third-party published - WeirdML: 39.5% (42.6: 42.6% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,357 (14.7: 7% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Vending-Bench 2: $202 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,219 (73.5: 37% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,408 (72.2: 36% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Grok 4.1 (xAI, United States) URL: https://opencharts.com/benchmarks/grok-4-1 Index: 43.1 (leave-one-out 20.6 to 61.3) · Grade: D · Rank: provisional · Verification: third-party published - Code Arena (WebDev): 1,211 (6.6: 3% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - APEX-Agents: 12.8% (27.0: 27.0% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Cybench: 39.0% (41.9: 41.9% of the best (93.0%)) — developer-reported; source: Epoch AI — AI Benchmarking Hub, https://data.x.ai/2025-11-17-grok-4-1-model-card.pdf - Text Arena: 1,465 (88.1: 44% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5 mini (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-mini Index: 43.0 (leave-one-out 37.0 to 48.4) · Grade: D · Rank: provisional · Verification: third-party published - GPQA Diamond: 75.0% (70.7: 70.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjP8EyLtPFsakgnsHe2oJM6.eval - Humanity's Last Exam: 19.4% (41.8: 41.8% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - ARC-AGI-2: 4.4% (4.7: 4.7% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 54.3% (55.2: 55.2% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 0.0% (0.0: 0.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 21.6% (28.6: 28.6% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SWE-bench Verified: 64.7% (77.5: 77.5% of the best (83.5%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F7nvdQy3AwKtoay2FBvfST5.eval - SciCode: 39.2% (63.2: 63.2% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 52.7% (56.7: 56.7% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Mock AIME 2024–2025: 86.7% (86.7: 86.7% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FR82gycJfm5jaCdh32tFkH4.eval - MATH Level 5: 97.8% (99.7: 99.7% of the best (98.1%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/math-level-5 - FrontierMath (Tiers 1–3): 46.7% (49.8: 49.8% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FYSR9WCR2SnJYvP7DJwX7YN.eval - FrontierMath Tier 4: 12.2% (12.5: 12.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F2WxRbjq5ZWTgiXZPqYUG4R.eval - ProofBench: 9.0% (9.0: 9.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 34.8% (41.1: 41.1% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - Vending-Bench 2: $-31 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,182 (64.0: 32% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,389 (67.4: 34% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text - Fiction.LiveBench: 62.5% (62.5: 62.5% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://fiction.live/stories/Fiction-liveBench-Mar-14-2025/oQdzQvKHw8JyXbN87/home ### Qwen3.6 35B-A3B (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-6-35b-a3b Index: 42.9 (leave-one-out 32.4 to 53.4) · Grade: D · Rank: provisional · Verification: third-party published - GPQA Diamond: 84.8% (84.6: 84.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F9kErpmgQ42wTvgJaFn478v.eval - CritPt: 0.3% (0.9: 0.9% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 35.8% (57.6: 57.6% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 34.5% (37.1: 37.1% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Mock AIME 2024–2025: 86.7% (86.7: 86.7% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F3n5KDcBaKqJh3WTg7qU2rs.eval - FrontierMath (Tiers 1–3): 20.4% (21.7: 21.7% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FYRvn36dmMyDbS5e4XkF4Xi.eval - Terminal-Bench: 23.0% (27.2: 27.2% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard ### MiniMax M2.1 (MiniMax, China) URL: https://opencharts.com/benchmarks/minimax-m2-1 Index: 42.1 (leave-one-out 30.3 to 54.6) · Grade: D · Rank: provisional · Verification: third-party published - Code Arena (WebDev): 1,387 (17.3: 9% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Terminal-Bench: 36.6% (43.2: 43.2% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - Text Arena: 1,384 (65.9: 33% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Grok 4.1 Fast (xAI, United States) URL: https://opencharts.com/benchmarks/grok-4-1-fast Index: 41.9 (leave-one-out 34.6 to 49.4) · Grade: D · Rank: provisional · Verification: third-party published - SimpleBench: 56.0% (68.4: 68.4% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - Code Arena (WebDev): 1,240 (7.8: 4% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 4.0% (4.0: 4.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vending-Bench 2: $1,107 (25.6: 25.6% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,195 (67.2: 34% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,430 (78.2: 39% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Grok 4.3 (xAI, United States) URL: https://opencharts.com/benchmarks/grok-4-3 Index: 41.3 (leave-one-out 23.9 to 51.3) · Grade: D · Rank: provisional · Verification: third-party published - GPQA Diamond: 88.8% (90.2: 90.2% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - CritPt: 8.0% (24.8: 24.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 33.2% (43.9: 43.9% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FY39ZRp553ULpUDBV2XrBbc.eval - SciCode: 47.3% (76.3: 76.3% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 49.9% (53.7: 53.7% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,356 (14.7: 7% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 93.3% (93.3: 93.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 42.8% (45.7: 45.7% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FbdsZjrG4ayzP8NAJg5q586.eval - FrontierMath Tier 4: 14.6% (15.0: 15.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F3iRQkf5v7Wu7YuZhxa76Gg.eval - ProofBench: 11.0% (11.0: 11.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vending-Bench 2: $35 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Vision Arena: 1,241 (79.6: 40% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,443 (81.7: 41% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.2 Codex (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-2-codex Index: 40.9 (leave-one-out 35.8 to 68.4) · Grade: D · Rank: provisional · Verification: third-party published - Code Arena (WebDev): 1,338 (13.3: 7% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Terminal-Bench: 66.5% (78.5: 78.5% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - APEX-Agents: 27.6% (58.2: 58.2% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents ### GPT-5.1 Codex Mini (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-1-codex-mini Index: 40.3 (leave-one-out 7.9 to 72.7) · Grade: D · Rank: provisional · Verification: third-party published - Code Arena (WebDev): 1,244 (7.9: 4% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Terminal-Bench: 61.6% (72.7: 72.7% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard ### GPT-5.4 nano (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-4-nano Index: 39.5 (leave-one-out 31.2 to 45.3) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 78.5% (75.6: 75.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FfZB3qGSohAVYzN97mniuUp.eval - ARC-AGI-2: 5.7% (6.0: 6.0% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 51.5% (52.3: 52.3% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 9.3% (28.6: 28.6% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 11.7% (15.5: 15.5% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F3QBsxKkbF2R2LEXTd6HDrJ.eval - SciCode: 46.9% (75.6: 75.6% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 49.2% (53.0: 53.0% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Mock AIME 2024–2025: 87.8% (87.8: 87.8% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjVv9m8DjkdpZaLPdUzAnG9.eval - FrontierMath (Tiers 1–3): 44.9% (47.9: 47.9% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FiKtopjTjyCWZqJaEjcc7ZL.eval - FrontierMath Tier 4: 12.2% (12.5: 12.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FSF6FYnnKztPCDWQc4v4EJa.eval - ProofBench: 5.0% (5.0: 5.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - APEX-Agents: 16.9% (35.7: 35.7% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vision Arena: 1,201 (68.9: 34% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,402 (70.7: 35% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.5 Flash (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-5-flash Index: 39.5 (leave-one-out 31.2 to 47.4) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 82.3% (81.0: 81.0% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FS7wbot7JR53BAjM8F9gaws.eval - SimpleQA Verified: 20.3% (26.9: 26.9% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FK8zTckSXLKLesoLBv24TWc.eval - Code Arena (WebDev): 1,238 (7.7: 4% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 84.4% (84.4: 84.4% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FawDKyKbPXRzBVZzfdQucef.eval - FrontierMath (Tiers 1–3): 18.2% (19.5: 19.5% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FafxoicJc4C6caVDwUpsiME.eval - Vending-Bench 2: $463 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,397 (69.3: 35% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-OSS 20B (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-oss-20b Index: 39.4 (leave-one-out 32.9 to 48.2) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 60.8% (50.6: 50.6% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FGEwyLqZT93Hp55fdtHrJhm.eval - CritPt: 1.4% (4.4: 4.4% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 34.4% (55.4: 55.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 40.9% (44.1: 44.1% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Mock AIME 2024–2025: 65.3% (65.3: 65.3% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FDB8ja6pEZWGocQbdciyisL.eval - Terminal-Bench: 3.4% (4.0: 4.0% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - Text Arena: 1,317 (50.2: 25% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Qwen3.5 9B (Alibaba, China) URL: https://opencharts.com/benchmarks/qwen3-5-9b Index: 38.9 (leave-one-out 29.5 to 48.3) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 79.0% (76.3: 76.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fo6migB5WP27uLr9nUWpCiM.eval - CritPt: 0.3% (0.9: 0.9% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 27.5% (44.4: 44.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Mock AIME 2024–2025: 61.7% (61.7: 61.7% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FYBMTDnW5LmCN6tBp2QYHzi.eval - Terminal-Bench: 9.2% (10.9: 10.9% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard ### GPT-5 nano (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-nano Index: 35.8 (leave-one-out 29.0 to 42.8) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 69.4% (62.8: 62.8% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F2pSDjMSGDh4HsAxkqUuNQP.eval - ARC-AGI-2: 2.6% (2.7: 2.7% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 20.7% (21.0: 21.0% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - SimpleQA Verified: 11.7% (15.5: 15.5% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - WeirdML: 38.1% (41.0: 41.0% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Mock AIME 2024–2025: 81.1% (81.1: 81.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FY9YnaJQWpPFAPhe6jtKMhH.eval - MATH Level 5: 95.2% (97.1: 97.1% of the best (98.1%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/math-level-5 - FrontierMath (Tiers 1–3): 20.0% (21.3: 21.3% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FJH3PjDrbrQndEcc8fh5Qre.eval - FrontierMath Tier 4: 2.4% (2.5: 2.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FP5v3FUnVVTbaqaCr43deKi.eval - ProofBench: 12.0% (12.0: 12.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 21.8% (25.7: 25.7% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - Vision Arena: 1,145 (55.1: 28% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,337 (54.6: 27% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text - Fiction.LiveBench: 21.9% (21.9: 21.9% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://fiction.live/stories/Fiction-liveBench-Mar-14-2025/oQdzQvKHw8JyXbN87/home ### MiniMax M2.7 (MiniMax, China) URL: https://opencharts.com/benchmarks/minimax-m2-7 Index: 35.4 (leave-one-out 25.7 to 43.8) · Grade: F · Rank: provisional · Verification: third-party published - CritPt: 0.6% (1.8: 1.8% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 47.0% (75.7: 75.7% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 37.0% (39.8: 39.8% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,398 (18.3: 9% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 3.0% (3.0: 3.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 45.1% (53.2: 53.2% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - Text Arena: 1,415 (74.2: 37% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GLM 4.6 (Zhipu AI, China) URL: https://opencharts.com/benchmarks/glm-4-6 Index: 34.2 (leave-one-out 20.0 to 44.4) · Grade: F · Rank: provisional · Verification: third-party published - CritPt: 1.1% (3.5: 3.5% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 38.4% (61.9: 61.9% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Code Arena (WebDev): 1,340 (13.5: 7% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Terminal-Bench: 24.5% (28.9: 28.9% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - APEX-Agents: 4.0% (8.4: 8.4% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Text Arena: 1,425 (76.7: 38% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 3.1 Flash-Lite (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-1-flash-lite Index: 34.1 (leave-one-out 34.1 to 58.8) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 81.8% (80.3: 80.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FnE4hmJo3QV6X5jeGWAeAze.eval - Humanity's Last Exam: 8.6% (18.6: 18.6% of the best (46.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://lastexam.ai - CritPt: 1.1% (3.5: 3.5% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 41.9% (67.5: 67.5% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 52.2% (56.2: 56.2% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,254 (8.4: 4% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 80.0% (80.0: 80.0% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FNMdaAXyqqPRmWfuxZUfHFF.eval - FrontierMath (Tiers 1–3): 27.7% (29.6: 29.6% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FMVZf59jt9uEoWG9w5W9apB.eval - APEX-Agents: 13.0% (27.4: 27.4% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - DeepResearch Bench: 37.3% (67.5: 67.5% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Vision Arena: 1,235 (77.9: 39% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,432 (78.8: 39% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Mistral Medium 3.5 (Mistral AI, France) URL: https://opencharts.com/benchmarks/mistral-medium-3-5 Index: 33.7 (leave-one-out 33.7 to 39.3) · Grade: F · Rank: provisional · Verification: third-party published - CritPt: 0.0% (0.0: 0.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 39.6% (63.8: 63.8% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 43.7% (47.1: 47.1% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - FrontierCode: 8.0% (15.0: 15.0% of the best (53.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://cognition.com/frontiercode - Code Arena (WebDev): 1,265 (8.9: 4% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 9.0% (9.0: 9.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,199 (68.3: 34% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,427 (77.3: 39% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.5 Instant (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-5-instant Index: 32.9 (leave-one-out 32.9 to 69.5) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 82.5% (81.3: 81.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/gpqa-diamond - CritPt: 0.0% (0.0: 0.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 48.6% (78.4: 78.4% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - Mock AIME 2024–2025: 68.1% (68.1: 68.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/otis-mock-aime-2024-2025 - FrontierMath (Tiers 1–3): 26.3% (28.1: 28.1% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FgGcHjE6ppy54ahN2pfxsGH.eval - FrontierMath Tier 4: 2.4% (2.5: 2.5% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F2UmuBfhBvYbbnTg5BnQqov.eval - Vision Arena: 1,278 (89.8: 45% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,474 (90.4: 45% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Claude Haiku 4.5 (Anthropic, United States) URL: https://opencharts.com/benchmarks/claude-haiku-4-5 Index: 32.7 (leave-one-out 26.7 to 37.6) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 71.2% (65.3: 65.3% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FGEaAdb9UmoUFbEtBx7WzAR.eval - ARC-AGI-2: 4.0% (4.2: 4.2% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 47.7% (48.4: 48.4% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 0.0% (0.0: 0.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SimpleQA Verified: 13.2% (17.5: 17.5% of the best (75.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/simpleqa-verified - SciCode: 43.3% (69.8: 69.8% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 45.4% (48.9: 48.9% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,329 (12.7: 6% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 66.7% (66.7: 66.7% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FneZvEzLi6yeL9z5KjGef7H.eval - MATH Level 5: 96.4% (98.2: 98.2% of the best (98.1%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/math-level-5 - Terminal-Bench: 35.5% (41.9: 41.9% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - APEX-Agents: 8.9% (18.8: 18.8% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $459 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - DeepResearch Bench: 45.5% (82.3: 82.3% of the best (55.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://deepresearch-bench.github.io - Text Arena: 1,413 (73.6: 37% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Gemini 3.5 Flash-Lite (Google, United States) URL: https://opencharts.com/benchmarks/gemini-3-5-flash-lite Index: 32.5 (leave-one-out 24.9 to 38.6) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 83.3% (82.4: 82.4% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fchbh6XstiqtUa5W9zuyfjS.eval - ARC-AGI-2: 10.3% (10.8: 10.8% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 53.5% (54.3: 54.3% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - CritPt: 0.0% (0.0: 0.0% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 40.9% (65.9: 65.9% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 39.0% (42.0: 42.0% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Code Arena (WebDev): 1,449 (23.8: 12% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Mock AIME 2024–2025: 71.1% (71.1: 71.1% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2Fmqsc4vFuKtAHXxbYyX9WwK.eval - FrontierMath (Tiers 1–3): 26.0% (27.7: 27.7% of the best (93.7%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F8o2J9uQLKDCTiCFwYdkJRK.eval - FrontierMath Tier 4: 0.0% (0.0: 0.0% of the best (97.6%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FjuNVfKBcYTwaKnk9GRA2zL.eval - ProofBench: 13.0% (13.0: 13.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Vision Arena: 1,266 (86.7: 43% expected win rate against the board leader (1,313)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/vision - Text Arena: 1,457 (85.6: 43% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### Nemotron 3 Super (NVIDIA, United States) URL: https://opencharts.com/benchmarks/nemotron-3-super Index: 29.6 (leave-one-out 25.3 to 49.5) · Grade: F · Rank: provisional · Verification: third-party published - CritPt: 3.1% (9.7: 9.7% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - SciCode: 36.0% (58.0: 58.0% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 38.0% (40.9: 40.9% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html ### MiniMax M2.5 (MiniMax, China) URL: https://opencharts.com/benchmarks/minimax-m2-5 Index: 28.9 (leave-one-out 19.3 to 35.2) · Grade: F · Rank: provisional · Verification: third-party published - ARC-AGI-2: 4.9% (5.1: 5.1% of the best (95.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - ARC-AGI-1: 63.7% (64.6: 64.6% of the best (98.5%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://arcprize.org/leaderboard - Code Arena (WebDev): 1,384 (17.0: 9% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 4.0% (4.0: 4.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 42.7% (50.4: 50.4% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - APEX-Agents: 6.2% (13.1: 13.1% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $-23 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,391 (67.6: 34% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-5.1 Codex (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-5-1-codex Index: 28.8 (leave-one-out 26.5 to 38.7) · Grade: F · Rank: provisional · Verification: third-party published - Code Arena (WebDev): 1,336 (13.1: 7% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - ProofBench: 9.0% (9.0: 9.0% of the best (100.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.vals.ai/benchmarks/proof_bench - Terminal-Bench: 60.4% (71.3: 71.3% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - METR Time Horizon: 3.7 h (77.8: 77.8% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ - APEX-Agents: 20.7% (43.7: 43.7% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents ### MiniMax M2 (MiniMax, China) URL: https://opencharts.com/benchmarks/minimax-m2 Index: 28.3 (leave-one-out 14.2 to 37.2) · Grade: F · Rank: provisional · Verification: third-party published - Code Arena (WebDev): 1,298 (10.7: 5% expected win rate against the board leader (1,797)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/webdev - Terminal-Bench: 30.0% (35.4: 35.4% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard - Vending-Bench 2: $161 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,346 (56.6: 28% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ### GPT-OSS 120B (OpenAI, United States) URL: https://opencharts.com/benchmarks/gpt-oss-120b Index: 27.8 (leave-one-out 21.5 to 34.1) · Grade: F · Rank: provisional · Verification: third-party published - GPQA Diamond: 75.8% (71.7: 71.7% of the best (95.8%), above chance) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2FLC5Rhog6wt2PT84CtJ483h.eval - SimpleBench: 22.1% (27.0: 27.0% of the best (81.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://simple-bench.com/ - CritPt: 1.1% (3.5: 3.5% of the best (32.3%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/critpt - Aider Polyglot: 41.8% (47.5: 47.5% of the best (88.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://aider.chat/docs/leaderboards/ - SciCode: 38.9% (62.7: 62.7% of the best (62.0%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://artificialanalysis.ai/evaluations/scicode - WeirdML: 48.2% (51.9: 51.9% of the best (92.9%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://htihle.github.io/weirdml.html - Mock AIME 2024–2025: 88.9% (88.9: 88.9% of the best (100.0%)) — run by Epoch AI; source: Epoch AI — AI Benchmarking Hub, https://logs.epoch.ai/inspect-viewer/36231d6d/viewer.html?log_file=https%3A%2F%2Flogs.epoch.ai%2Finspect_ai_logs%2F3QjP5HPi8tKUzHtNJktfv4.eval - Terminal-Bench: 18.7% (22.1: 22.1% of the best (84.7%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://www.tbench.ai/leaderboard/terminal-bench/2.0 - METR Time Horizon: 42 min (53.8: 53.8% of the best (17.4 h) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ - APEX-Agents: 4.7% (9.9: 9.9% of the best (47.4%)) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://epoch.ai/benchmarks/apex-agents - Vending-Bench 2: $-22 (0.0: 0.0% of the best ($11,182) on a log scale) — official leaderboard or third-party evaluator; source: Epoch AI — AI Benchmarking Hub, https://andonlabs.com/evals/vending-bench-2 - Text Arena: 1,352 (58.2: 29% expected win rate against the board leader (1,507)) — official leaderboard or third-party evaluator; source: Arena (LMArena) — Leaderboard Dataset, https://arena.ai/leaderboard/text ## Benchmarks covered - GPQA Diamond (Reasoning; Epoch AI (internal runs); run by Epoch AI; 198 items; Caveat: Frontier models cluster in the eighties and nineties; Epoch's standard errors on this test are 2 to 3 points.) — 198 PhD-level four-option questions in physics, chemistry and biology, written to be Google-proof. Epoch AI runs most models itself 16 times; random guessing scores 25%. https://opencharts.com/benchmarks/tests/gpqa-diamond - Humanity's Last Exam (Reasoning; Center for AI Safety & Scale AI; official leaderboard or third-party evaluator; 2500 items) — 2,500 expert-written questions across a hundred fields at the frontier of human knowledge. Accuracy on the full exam, from the exam's own leaderboard. https://opencharts.com/benchmarks/tests/hle - ARC-AGI-2 (Reasoning; ARC Prize Foundation; official leaderboard or third-party evaluator; 120 items; Caveat: Scores depend on the compute budget a lab chose; ARC Prize publishes cost per task beside every score and this ranking does not.) — Abstract visual reasoning puzzles that are easy for people and designed to resist memorization — the benchmark built to measure general fluid intelligence. Semi-private evaluation set, verified by ARC Prize. https://opencharts.com/benchmarks/tests/arc-agi-2 - ARC-AGI-1 (Reasoning; ARC Prize Foundation; official leaderboard or third-party evaluator; 100 items; Caveat: Close to saturated at the frontier; scores depend on the compute budget a lab chose.) — The original Abstraction and Reasoning Corpus — grid puzzles solved from a handful of examples. Semi-private evaluation set, verified by ARC Prize. https://opencharts.com/benchmarks/tests/arc-agi - LiveBench (Reasoning; LiveBench; official leaderboard or third-party evaluator; Caveat: Epoch's copy of this leaderboard lags the live site, so few current models have a score here.) — A contamination-resistant benchmark refreshed monthly — reasoning, coding, math, data analysis, language and instruction following. Global average. https://opencharts.com/benchmarks/tests/livebench - SimpleBench (Reasoning; SimpleBench / LM Council; official leaderboard or third-party evaluator) — Trick questions about the everyday world where humans score in the eighties — spatio-temporal reasoning, social intelligence and adversarial phrasing. https://opencharts.com/benchmarks/tests/simplebench - CritPt (Reasoning; CritPt / Artificial Analysis; official leaderboard or third-party evaluator) — Unpublished research-level physics problems graded by an official server — accuracy on the full set, as run by Artificial Analysis. https://opencharts.com/benchmarks/tests/critpt - SimpleQA Verified (Knowledge; Epoch AI (internal runs); run by Epoch AI; 1000 items) — A thousand verified short-answer factual questions — how often the model is right, without being allowed to search. Epoch AI runs it. https://opencharts.com/benchmarks/tests/simpleqa-verified - SWE-bench Verified (Coding; Epoch AI (internal runs); run by Epoch AI; 500 items; Caveat: One shared scaffold for every model; labs' own numbers with bespoke scaffolds can run 10 points higher and are not used here.) — 500 human-validated GitHub issues the model must resolve in the real repository so the hidden tests pass. Epoch AI runs every model in one standardized scaffold. https://opencharts.com/benchmarks/tests/swe-bench-verified - Aider Polyglot (Coding; Aider; official leaderboard or third-party evaluator; 225 items) — 225 of the hardest Exercism exercises across six languages, edited in place and checked by the tests. https://opencharts.com/benchmarks/tests/aider-polyglot - SciCode (Coding; SciCode / Artificial Analysis; official leaderboard or third-party evaluator; 338 items) — Research-grade scientific programming: 338 sub-problems from 80 research problems across physics, chemistry, biology, materials and math that must pass unit tests. As run by Artificial Analysis. https://opencharts.com/benchmarks/tests/scicode - WeirdML (Coding; Håvard Ihle; official leaderboard or third-party evaluator) — Unusual machine-learning tasks the model must solve end to end by writing and running training code. https://opencharts.com/benchmarks/tests/weirdml - FrontierCode (Coding; Cognition; official leaderboard or third-party evaluator; Caveat: Published by Cognition, a coding-agent vendor, from its own harness; not independently reproduced.) — Hard, realistic software tasks run through a coding-agent harness. Main score. https://opencharts.com/benchmarks/tests/frontiercode - CursorBench (Coding; Cursor; official leaderboard or third-party evaluator; Caveat: Published by Cursor from its own product traffic and harness; not independently reproduced.) — Real coding-agent requests drawn from everyday editor sessions, scored against the change the developer actually shipped. https://opencharts.com/benchmarks/tests/cursorbench - Code Arena (WebDev) (Coding; Arena; official leaderboard or third-party evaluator) — Blind head-to-head votes on web apps two models built from the same prompt. Arena score. https://opencharts.com/benchmarks/tests/arena-webdev - Mock AIME 2024–2025 (Math; Epoch AI (internal runs); run by Epoch AI; 45 items; Caveat: 45 problems: one problem is 2.2 points, so gaps of a few points are within noise.) — Competition mathematics from the OTIS Mock AIME sets — 45 problems, exact-answer graded, run 16 times by Epoch AI. https://opencharts.com/benchmarks/tests/aime-2024-2025 - MATH Level 5 (Math; Epoch AI (internal runs); run by Epoch AI; Caveat: Epoch flags likely training contamination on this set; treat small gaps as meaningless.) — The hardest tier of the MATH competition dataset. Saturated at the frontier; useful for the long tail. https://opencharts.com/benchmarks/tests/math-level-5 - FrontierMath (Tiers 1–3) (Math; Epoch AI (internal runs); run by Epoch AI; 290 items; Caveat: Commissioned by OpenAI, which has access to much of the problem set; Epoch discloses this and keeps a holdout it does not share.) — 290 unpublished research-level math problems written by professional mathematicians, tiers 1–3. Epoch AI runs it. https://opencharts.com/benchmarks/tests/frontiermath - FrontierMath Tier 4 (Math; Epoch AI (internal runs); run by Epoch AI; 48 items; Caveat: 48 problems: one problem is about 2 points. Same OpenAI commissioning disclosure as tiers 1 to 3.) — The hardest FrontierMath tier — 48 problems that take expert mathematicians days. Epoch AI runs it. https://opencharts.com/benchmarks/tests/frontiermath-tier-4 - ProofBench (Math; Vals AI; official leaderboard or third-party evaluator) — Full written proofs, not final answers, graded for rigor. https://opencharts.com/benchmarks/tests/proofbench - Terminal-Bench (Agentic; Terminal-Bench; official leaderboard or third-party evaluator; Caveat: Harnesses differ by model, so this compares model-plus-harness systems, not models alone.) — Real tasks done in a terminal — set up servers, fix builds, wrangle data — checked by tests. Best published agent harness per model. https://opencharts.com/benchmarks/tests/terminal-bench - GDPval (Agentic; OpenAI (external evaluations); official leaderboard or third-party evaluator; 220 items; Caveat: Authored and graded by OpenAI, which also competes on it. Treat as a lab-run leaderboard.) — Economically valuable work across 44 occupations, judged by industry professionals against deliverables from their peers. Win rate on the 220-task gold set. https://opencharts.com/benchmarks/tests/gdpval - METR Time Horizon (Agentic; METR; official leaderboard or third-party evaluator; Caveat: METR publishes wide confidence intervals around each horizon; the point estimate is used here.) — The length of software task (in human-expert minutes) a model completes with 50% reliability. Shown in minutes; normalized on a log scale from one minute to the frontier. https://opencharts.com/benchmarks/tests/metr-time-horizons - OSWorld 2.0 (Agentic; XLANG Lab; official leaderboard or third-party evaluator; 108 items; Caveat: A separate, far harder benchmark from OSWorld-Verified: binary completion is low for every model, and the leaderboard also reports a partial-credit score.) — 108 long-horizon computer-use workflows across real desktop and web applications; a skilled person takes a median of about 1.6 hours per task. Share of tasks completed in full at the 500-step budget. https://opencharts.com/benchmarks/tests/osworld-2 - APEX-Agents (Agentic; Mercor; official leaderboard or third-party evaluator) — Professional-services tasks — consulting, law, finance — completed as an agent and graded against expert rubrics. Pass@1. https://opencharts.com/benchmarks/tests/apex-agents - Vending-Bench 2 (Agentic; Andon Labs; official leaderboard or third-party evaluator; Caveat: Andon Labs estimates a strong human operator at roughly $63,000, so every model is far from the ceiling.) — Run a simulated vending business for a year from a $500 float — ordering, pricing, cash flow. Mean final balance over five runs; normalized on a log scale from the starting balance to the frontier. https://opencharts.com/benchmarks/tests/vending-bench-2 - Cybench (Agentic; Stanford; developer-reported; 40 items; Caveat: Developer-reported: the numbers come from the labs' own model cards, collected by Epoch AI, not from an independent run.) — 40 professional capture-the-flag cybersecurity challenges solved unguided. Share solved. https://opencharts.com/benchmarks/tests/cybench - DeepResearch Bench (Agentic; DeepResearch Bench; official leaderboard or third-party evaluator; 100 items) — PhD-level research briefs written by an agent with web access, graded on coverage, insight and citations. https://opencharts.com/benchmarks/tests/deep-research-bench - Vision Arena (Multimodal; Arena; official leaderboard or third-party evaluator) — Blind head-to-head votes on prompts that include an image. Style-controlled Arena score. https://opencharts.com/benchmarks/tests/arena-vision - Text Arena (Human preference; Arena; official leaderboard or third-party evaluator) — Millions of blind head-to-head votes from real users. Style-controlled Arena score — the market's human-preference standard. https://opencharts.com/benchmarks/tests/arena-text - Fiction.LiveBench (Long context; Fiction.live; official leaderboard or third-party evaluator) — Deep comprehension of long stories: questions that need the whole text, at a 120K-token context length. https://opencharts.com/benchmarks/tests/fiction-livebench ## Attribution - Epoch AI, 'AI Benchmarking Hub'. Published online at epoch.ai. Retrieved from https://epoch.ai/benchmarks. License: CC BY 4.0. - Arena Leaderboard Dataset, lmarena-ai (Hugging Face, CC BY 4.0). Live boards at https://arena.ai/leaderboard. License: CC BY 4.0. - Internal scores: OpenCharts internal evaluation — protocol at https://opencharts.com/benchmarks/methodology#internal.