All benchmarks
Agentic

Vending-Bench 2

Run a simulated vending business for a year from a $500 float — ordering, pricing, cash flow. Mean final balance over five runs; normalized on a log scale from the starting balance to the frontier.

Long-horizon tasks with tools: terminals, desktops, browsers, whole jobs. One of 8 comparable tests in agentic, a category the Index averages.

Read with care. Andon Labs estimates a strong human operator at roughly $63,000, so every model is far from the ceiling.

Publisher: Andon Labs

What a model is asked to do

Run a simulated vending business for a year: ordering, pricing and cash flow. The score is the final balance.

For exampleStart with a small cash float and one machine; negotiate supplier prices, set a menu and keep it stocked through the seasons.

Why it matters. Long-horizon decisions whose consequences compound. It exposes models that drift or forget.

The example is original and illustrative, not an item from the dataset.

Models scored here
53
Who produced the numbers
Official leaderboard
Best published result
$11,182
Log-scale floor
$500
License
CC BY 4.0 (via Epoch AI)
Direction
Higher is better
Scale
An open-ended value. For the Index, each score is placed on a log scale from a fixed floor ($500) to the best published result ($11,182), so the frontier reads 100 and doubling counts the same anywhere on the scale.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

53 models on Vending-Bench 2

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100, on a log scale). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.

#ModelScoreProduced by
1Claude Opus 5Anthropic$11,182Leaderboard
2Claude Opus 4.7Anthropic$10,937Leaderboard
3GPT-5.6 SolOpenAI$9,619Leaderboard
4Grok 4.6xAI$9,047Leaderboard
5GLM 5.2Zhipu AI$8,314Leaderboard
6GLM 5.3Zhipu AI$8,164Leaderboard
7Claude Opus 4.6Anthropic$8,018Leaderboard
8GPT-5.5OpenAI$7,524Leaderboard
9GPT-5.6 TerraOpenAI$7,343Leaderboard
10Claude Sonnet 4.6Anthropic$7,204Leaderboard
11Muse Spark 1.1Meta$6,520Leaderboard
12Claude Sonnet 5Anthropic$6,378Leaderboard
13Kimi K2.6Moonshot AI$6,205Leaderboard
14GPT-5.4OpenAI$6,144Leaderboard
15GPT-5.3 CodexOpenAI$5,940Leaderboard
16Claude Opus 4.8Anthropic$5,787Leaderboard
17Claude Fable 5Anthropic$5,680Leaderboard
18GLM 5.1Zhipu AI$5,634Leaderboard
19Gemini 3 ProGoogle$5,478Leaderboard
20Claude Fable 5.1Anthropic$5,422Leaderboard
21Gemini 3.5 FlashGoogle$5,396Leaderboard
22Kimi K3Moonshot AI$5,165Leaderboard
23Qwen3.6 PlusAlibaba$5,115Leaderboard
24Kimi K2.7 CodeMoonshot AI$5,083Leaderboard
25Claude Opus 4.5Anthropic$4,967Leaderboard
26Grok 4.20xAI$4,663Leaderboard
27GLM 5Zhipu AI$4,432Leaderboard
28Qwen3.6 MaxAlibaba$4,254Leaderboard
29GPT-5.6 LunaOpenAI$4,095Leaderboard
30Grok 4.5xAI$3,887Leaderboard
31Claude Sonnet 4.5Anthropic$3,839Leaderboard
32Gemini 3.1 ProGoogle$3,774Leaderboard
33Gemini 3 FlashGoogle$3,635Leaderboard
34GPT-5.2OpenAI$3,591Leaderboard
35DeepSeek V4 ProDeepSeek$3,285Leaderboard
36GLM 4.7Zhipu AI$2,377Leaderboard
37MiniMax M3MiniMax$2,158Leaderboard
38GPT-5.1OpenAI$1,473Leaderboard
39Kimi K2.5Moonshot AI$1,198Leaderboard
40Grok 4.1 FastxAI$1,107Leaderboard
41DeepSeek V3.2DeepSeek$1,034Leaderboard
42Gemini 2.5 ProGoogle$574Leaderboard
43Gemini 2.5 FlashGoogle$549Leaderboard
44Qwen3 MaxAlibaba$72Leaderboard
45Qwen3.5 PlusAlibaba$1Leaderboard
46Qwen3.5 27BAlibaba$202Leaderboard
47GPT-5 miniOpenAI$-31Leaderboard
48Grok 4.3xAI$35Leaderboard
49Qwen3.5 FlashAlibaba$463Leaderboard
50Claude Haiku 4.5Anthropic$459Leaderboard
51MiniMax M2.5MiniMax$-23Leaderboard
52MiniMax M2MiniMax$161Leaderboard
53GPT-OSS 120BOpenAI$-22Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.