Aider polyglot — Epoch AI
Aider polyglot — Epoch AI is run by Epoch AI. It has ranked 72 models, of which the catalogue holds 46, scoring from 3.6 to 88.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 72 | GPT-5 | Percent correct (high) | 88 |
| 3rdof 72 | o3-pro | Percent correct (high) | 84.9 |
| 4thof 72 | Gemini 2.5 Pro Preview 06-05 | Percent correct | 83.1 |
| 6thof 72 | o3 | Percent correct (high) | 81.3 |
| 11thof 72 | Gemini 2.5 Pro Preview 05-06 | Percent correct | 76.9 |
| 14thof 72 | deepseek-reasoner | Percent correct | 74.2 |
| 16thof 72 | Gemini 2.5 Pro | Percent correct | 72.9 |
| 17thof 72 | Claude Opus 4 | Percent correct | 72 |
| 18thof 72 | o4-mini | Percent correct (high) | 72 |
| 19thof 72 | DeepSeek-R1 (0528) | Percent correct | 71.4 |
| 21stof 72 | DeepSeek V3.2 Exp | Percent correct | 70.2 |
| 22ndof 72 | deepseek-chat | Percent correct | 70.2 |
| 23rdof 72 | Claude 3.7 Sonnet | Percent correct | 64.9 |
| 24thof 72 | o1 | Percent correct (high) | 61.7 |
| 25thof 72 | Claude Sonnet 4 | Percent correct | 61.3 |
| 27thof 72 | o3-mini | Percent correct (high) | 60.4 |
| 28thof 72 | Qwen3 235B A22B | Percent correct | 59.6 |
| 29thof 72 | Qwen3-235B-A22B-Instruct-2507 | Percent correct | 59.6 |
| 30thof 72 | Kimi K2 Instruct | Percent correct | 59.1 |
| 31stof 72 | Kimi K2 0905 | Percent correct | 59.1 |
| 32ndof 72 | DeepSeek-R1 | Percent correct | 56.9 |
| 34thof 72 | Gemini 2.5 Flash | Percent correct | 55.1 |
| 35thof 72 | DeepSeek-V3 0324 | Percent correct | 55.1 |
| 37thof 72 | Grok 3 | Percent correct | 53.3 |
| 38thof 72 | GPT-4.1 fine-tuned | Percent correct | 52.4 |
| 39thof 72 | Claude 3.5 Sonnet | Percent correct | 51.6 |
| 40thof 72 | Grok 3 Mini | Percent correct (high) | 49.3 |
| 41stof 72 | DeepSeek-V3 | Percent correct | 48.4 |
| 46thof 72 | GPT OSS 120B | Percent correct (high) | 41.8 |
| 48thof 72 | Qwen3 32B | Percent correct | 40 |
| 52ndof 72 | o1-mini | Percent correct | 32.9 |
| 53rdof 72 | GPT-4.1 mini | Percent correct | 32.4 |
| 54thof 72 | Claude Haiku 3.5 | Percent correct | 28 |
| 56thof 72 | GPT-4o (2024-08-06) | Percent correct | 23.1 |
| 57thof 72 | Gemini 2.0 Flash | Percent correct | 22.2 |
| 59thof 72 | QwQ-32B | Percent correct | 20.9 |
| 61stof 72 | GPT-4o (2024-11-20) | Percent correct | 18.2 |
| 62ndof 72 | DeepSeek-V2.5 | Percent correct | 17.8 |
| 63rdof 72 | Qwen2.5 Coder 32B Instruct | Percent correct | 16.4 |
| 64thof 72 | Llama-4-Maverick-17B-128E-Instruct | Percent correct | 15.6 |
| 66thof 72 | c4ai-command-a-03-2025 | Percent correct | 12 |
| 67thof 72 | Codestral 25.01 | Percent correct | 11.1 |
| 68thof 72 | openhands-lm-32b-v0.1 | Percent correct | 10.2 |
| 69thof 72 | GPT-4.1 nano | Percent correct | 8.9 |
| 71stof 72 | Gemma 3 27B | Percent correct | 4.9 |
| 72ndof 72 | GPT-4o-mini (2024-07-18) | Percent correct | 3.6 |
Read from the board on 2026-09-16