Fiction.LiveBench
Deep comprehension of long stories: questions that need the whole text, at a 120K-token context length.
Keeping track of a story across a very long document. Long context is a single-test category in this snapshot (fewer than 2 comparable tests), so it is shown beside the Index and never averaged into it.
Publisher: Fiction.liveWhat a model is asked to do
Answer questions that need the whole of a very long story, at a 120K-token context length.
For exampleHalfway through the novel, which promise from chapter two does the narrator break, and who is the first to notice?
Why it matters. Long documents are the normal case at work. This shows whether a model still holds the thread at the end.
The example is original and illustrative, not an item from the dataset.
- Models scored here
- 7
- Who produced the numbers
- Official leaderboard
- Best published result
- 100.0%
- License
- CC BY 4.0 (via Epoch AI)
- Direction
- Higher is better
- Scale
- Published as a share of items solved. For the Index, each score is a share of the best published result (100.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
- Provenance
- Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.
Snapshot September 8, 2026
Ranking on this test
7 models on Fiction.LiveBench
Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page.
| # | Model | Score | vs frontier | Run | Produced by |
|---|---|---|---|---|---|
| 1 | GPT-5OpenAI | 96.9% | 96.9 | 2025-08-07 · medium | Leaderboard |
| 2 | Grok 4xAI | 96.9% | 96.9 | — | Leaderboard |
| 3 | Kimi K2.5Moonshot AI | 78.1% | 78.1 | — | Leaderboard |
| 4 | Grok 4 FastxAI | 75.0% | 75.0 | — | Leaderboard |
| 5 | GPT-5 miniOpenAI | 62.5% | 62.5 | 2025-08-07 · medium | Leaderboard |
| 6 | Kimi K2Moonshot AI | 40.6% | 40.6 | — | Leaderboard |
| 7 | GPT-5 nanoOpenAI | 21.9% | 21.9 | 2025-08-07 · medium | Leaderboard |

The best models, ranked here, working inside Theo.
28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.