All benchmarks
Coding

Aider Polyglot

225 of the hardest Exercism exercises across six languages, edited in place and checked by the tests.

Writing, fixing and shipping real code, judged by tests or users. One of 7 comparable tests in coding, a category the Index averages.

Publisher: Aider

What a model is asked to do

Edit code in place to pass the tests of hard exercises across six programming languages.

For exampleImplement a sliding-window rate limiter in Rust so the provided test suite passes. Edit the existing files; do not rewrite them.

Why it matters. Precise, multi-language edits are what a coding assistant does all day inside an editor.

The example is original and illustrative, not an item from the dataset.

Models scored here
5
Who produced the numbers
Official leaderboard
Items graded
n = 225 (0.4% each)
Best published result
88.0%
License
Apache 2.0 (via Epoch AI)
Direction
Higher is better
Scale
Published as a share of items solved. For the Index, each score is a share of the best published result (88.0%), so the frontier reads 100. Guessing earns nothing on this test, so no chance correction applies.
Provenance
Produced by the benchmark's own leaderboard or a third-party evaluator, not by the model's maker.

Snapshot September 8, 2026

Ranking on this test

5 models on Aider Polyglot

Best published run per model, and how close each one comes to the best published result on this test (the frontier reads 100). Every number links to the publisher that produced it. Internal runs are marked and carry their run details on the model page. One item on this test is 0.4 points, so read gaps smaller than that as noise.

#ModelScoreProduced by
1GPT-5OpenAI88.0%Leaderboard
2Grok 4xAI79.6%Leaderboard
3DeepSeek V3.2DeepSeek74.2%Leaderboard
4Kimi K2Moonshot AI59.1%Leaderboard
5GPT-OSS 120BOpenAI41.8%Leaderboard
Theo

The best models, ranked here, working inside Theo.

28 of the models on this page run inside Theo today. Theo picks the right one for each step and always shows which engine answered.