Dtbench — Epoch AI
Dtbench — Epoch AI is run by Epoch AI. It has ranked 161 models, of which the catalogue holds 136, scoring from 41.65 to 98.4.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 161 | Claude Fable 5 | Accuracy (max) | 98.4 |
| 2ndof 161 | Claude Opus 5 | Accuracy (max) | 97.6 |
| 3rdof 161 | Grok 4.6 | Accuracy (xhigh) | 97.33 |
| 4thof 161 | Gemini 3.7 Flash | Accuracy (high) | 96.8 |
| 5thof 161 | Grok 4.5 | Accuracy (high) | 96.53 |
| 7thof 161 | GPT-5.5 | Accuracy (xhigh) | 96 |
| 8thof 161 | GPT-5.5 Pro | Accuracy (xhigh) | 96 |
| 9thof 161 | GPT-5.6 Sol | Accuracy (max) | 95.47 |
| 10thof 161 | Gemini 3.1 Pro Preview | Accuracy (high) | 95.47 |
| 11thof 161 | Gemini 3.6 Flash | Accuracy (high) | 95.47 |
| 12thof 161 | Claude Opus 4.8 | Accuracy (max) | 94.93 |
| 13thof 161 | Muse Spark 1.2 | Accuracy (xhigh) | 94.67 |
| 14thof 161 | Gemini 3.5 Flash | Accuracy (high) | 94.67 |
| 15thof 161 | Claude Opus 4.7 | Accuracy (max) | 94.65 |
| 16thof 161 | GPT-5.4 | Accuracy (xhigh) | 94.4 |
| 17thof 161 | Muse Spark 1.1 | Accuracy (high) | 94.4 |
| 18thof 161 | GLM 5.2 | Accuracy (max) | 93.6 |
| 19thof 161 | GPT-5.6 Terra | Accuracy (max) | 92.53 |
| 20thof 161 | Claude Sonnet 5 | Accuracy (max) | 92.48 |
| 21stof 161 | Qwen3.7 Max | Accuracy (max) | 92.27 |
| 22ndof 161 | Qwen 3.8 Max | Accuracy (xhigh) | 92 |
| 23rdof 161 | Claude Opus 4.6 | Accuracy (max) | 91.2 |
| 24thof 161 | Kimi K3 | Accuracy (max) | 91.2 |
| 25thof 161 | GPT-5.2 | Accuracy (xhigh) | 90.93 |
| 26thof 161 | DeepSeek V4 Flash (0731) | Accuracy (max) | 90.93 |
| 27thof 161 | Kimi K2.6 | Accuracy | 90.93 |
| 28thof 161 | DeepSeek V4 Pro | Accuracy (max) | 90.67 |
| 29thof 161 | GPT-5 | Accuracy (high) | 90.67 |
| 30thof 161 | Grok 4.3 | Accuracy (high) | 90.67 |
| 31stof 161 | GPT-5.1 | Accuracy (high) | 90.13 |
| 32ndof 161 | Grok 4.20 | Accuracy | 90.13 |
| 33rdof 161 | Nemotron 3 Ultra | Accuracy | 90.13 |
| 34thof 161 | Claude Sonnet 4.6 | Accuracy (max) | 89.87 |
| 35thof 161 | Claude Opus 4.5 | Accuracy (high) | 89.87 |
| 36thof 161 | GPT-5.6 Luna | Accuracy (max) | 89.07 |
| 37thof 161 | Gemini 3 Flash | Accuracy (high) | 89.07 |
| 38thof 161 | Grok 4.1 Fast Reasoning | Accuracy | 87.73 |
| 40thof 161 | Inkling | Accuracy (xhigh) | 87.47 |
| 41stof 161 | Qwen3.5 397B A17B | Accuracy | 87.47 |
| 42ndof 161 | Qwen3.6 Max Preview | Accuracy (max) | 87.2 |
| 43rdof 161 | o3-pro | Accuracy (high) | 86.93 |
| 44thof 161 | DeepSeek V4 Flash | Accuracy (max) | 86.4 |
| 45thof 161 | DeepSeek v3.2 | Accuracy | 85.6 |
| 46thof 161 | o3 | Accuracy (high) | 84.8 |
| 47thof 161 | MiMo-V2.5-Pro | Accuracy | 84.53 |
| 48thof 161 | Qwen3.5-122B-A10B | Accuracy (none) | 84.27 |
| 49thof 161 | Qwen 3.7 Plus | Accuracy | 84 |
| 50thof 161 | Gemini 3.5 Flash-Lite | Accuracy (high) | 83.47 |
| 51stof 161 | Claude Sonnet 4.5 | Accuracy | 83.2 |
| 52ndof 161 | Qwen3.5-Flash | Accuracy | 82.93 |
| 53rdof 161 | Grok 4 Fast | Accuracy | 82.67 |
| 54thof 161 | Gemma 4 31B | Accuracy | 82.67 |
| 56thof 161 | Gemini 2.5 Pro | Accuracy | 82.4 |
| 57thof 161 | Qwen3.5-27B | Accuracy | 82.4 |
| 59thof 161 | Qwen3.6 Plus | Accuracy | 81.87 |
| 60thof 161 | Claude Opus 4 | Accuracy | 81.6 |
| 61stof 161 | DeepSeek V3.1 Terminus | Accuracy | 81.33 |
| 62ndof 161 | Qwen-Plus | Accuracy | 81.07 |
| 63rdof 161 | GPT-5 mini | Accuracy (high) | 80.53 |
| 64thof 161 | Qwen 3.5 Plus | Accuracy | 80.53 |
| 65thof 161 | GPT-5.4 nano | Accuracy (xhigh) | 80.27 |
| 66thof 161 | Qwen3 235B A22B Thinking 2507 | Accuracy | 80.27 |
| 67thof 161 | GPT-5.4 mini | Accuracy (xhigh) | 80 |
| 68thof 161 | Claude Opus 4.1 | Accuracy | 80 |
| 69thof 161 | Qwen3.5-35B-A3B | Accuracy | 80 |
| 70thof 161 | MiniMax M3 | Accuracy | 78.93 |
| 71stof 161 | Qwen3-235B-A22B-Instruct-2507 | Accuracy | 78.4 |
| 72ndof 161 | Qwen3.6 27B | Accuracy | 78.13 |
| 73rdof 161 | o4-mini | Accuracy (high) | 77.6 |
| 74thof 161 | Qwen3.6-Flash | Accuracy | 77.07 |
| 75thof 161 | Claude Sonnet 4 | Accuracy | 77.07 |
| 76thof 161 | Gemini 3.1 Flash-Lite | Accuracy (high) | 76.8 |
| 77thof 161 | Gemini 2.5 Flash | Accuracy | 76.53 |
| 78thof 161 | GPT OSS 120B | Accuracy (high) | 76.27 |
| 79thof 161 | Qwen3 235B A22B | Accuracy | 75.73 |
| 81stof 161 | Gemma 4 26B A4B | Accuracy | 74.93 |
| 82ndof 161 | o1 | Accuracy (high) | 74.67 |
| 83rdof 161 | Qwen3.6 35B A3B | Accuracy | 73.87 |
| 84thof 161 | Claude Haiku 4.5 | Accuracy | 73.6 |
| 85thof 161 | Qwen3.5 9B | Accuracy | 71.2 |
| 86thof 161 | Mistral Small 4 | Accuracy | 70.93 |
| 87thof 161 | Qwen3 30B A3B Thinking 2507 | Accuracy | 69.33 |
| 88thof 161 | o3-mini | Accuracy (high) | 68.8 |
| 89thof 161 | GPT-4.1 mini | Accuracy | 68.8 |
| 90thof 161 | GPT-4.1 fine-tuned | Accuracy | 68.27 |
| 91stof 161 | GPT OSS 20B | Accuracy (high) | 68 |
| 92ndof 161 | Claude 3.5 Sonnet | Accuracy | 67.84 |
| 94thof 161 | Qwen3 32B | Accuracy | 67.47 |
| 96thof 161 | Qwen3 30B A3B Instruct 2507 | Accuracy | 67.2 |
| 98thof 161 | Mistral Large 3 | Accuracy | 65.07 |
| 99thof 161 | DeepSeek-V3 0324 | Accuracy | 64.8 |
| 100thof 161 | GPT-4o (2024-05-13) | Accuracy | 64.53 |
| 101stof 161 | Qwen3 14B | Accuracy | 64 |
| 102ndof 161 | gemini-2.0-flash-001 | Accuracy | 63.2 |
| 103rdof 161 | Qwen2.5 72B Instruct | Accuracy | 62.93 |
| 105thof 161 | GPT-4 (0613) | Accuracy | 62.67 |
| 106thof 161 | deepseek-chat | Accuracy | 62.67 |
| 107thof 161 | GPT-5 nano | Accuracy (high) | 62.67 |
| 108thof 161 | Mistral Medium 2505 | Accuracy | 62.27 |
| 109thof 161 | Llama-4-Maverick-17B-128E-Instruct | Accuracy | 61.87 |
| 110thof 161 | Claude 3 Opus | Accuracy | 61.6 |
| 111thof 161 | GPT-4 Turbo | Accuracy | 61.6 |
| 112thof 161 | Meta-Llama-3.1-405B-Instruct | Accuracy | 61.38 |
| 113thof 161 | c4ai-command-a-03-2025 | Accuracy | 61.33 |
| 115thof 161 | Mistral Large 2407 | Accuracy | 61.21 |
| 117thof 161 | Qwen3 30B A3B | Accuracy | 60.27 |
| 118thof 161 | Meta-Llama-3.1-70B-Instruct | Accuracy | 60 |
| 120thof 161 | Qwen3 8B | Accuracy | 59.73 |
| 121stof 161 | Llama 3.3 70B Instruct | Accuracy | 59.47 |
| 123rdof 161 | mistral-small-2503 | Accuracy | 58.64 |
| 125thof 161 | Claude Haiku 3.5 | Accuracy | 56.66 |
| 127thof 161 | Mistral Large (24.02) | Accuracy | 55.9 |
| 128thof 161 | Mixtral-8x22B-Instruct-v0.1 | Accuracy | 55.14 |
| 129thof 161 | c4ai-command-r-plus-08-2024 | Accuracy | 54.93 |
| 130thof 161 | GPT-4o-mini (2024-07-18) | Accuracy | 54.4 |
| 131stof 161 | Meta-Llama-3-70B-Instruct | Accuracy | 54.16 |
| 133rdof 161 | Claude 3 Sonnet | Accuracy | 53.6 |
| 135thof 161 | Mistral Small (24.02) | Accuracy | 52.96 |
| 136thof 161 | Gemini 2.0 Flash-Lite | Accuracy | 52.53 |
| 137thof 161 | Gemma 3 27B | Accuracy | 52.53 |
| 138thof 161 | GPT-4.1 nano | Accuracy | 52.53 |
| 139thof 161 | Claude 2.0 | Accuracy | 51.85 |
| 142ndof 161 | Claude 2.1 | Accuracy | 50.95 |
| 143rdof 161 | Llama 3.1 8B Instruct | Accuracy | 50.93 |
| 144thof 161 | Gemma 3 4B | Accuracy | 50.93 |
| 145thof 161 | Claude 3 Haiku | Accuracy | 50.13 |
| 149thof 161 | Gemma 3 12B | Accuracy | 48.8 |
| 150thof 161 | Mistral Nemo | Accuracy | 48.61 |
| 151stof 161 | GPT-3.5 Turbo | Accuracy | 48.53 |
| 152ndof 161 | Gemma 2 27B | Accuracy | 48 |
| 153rdof 161 | Qwen2.5 7B Instruct | Accuracy | 47.73 |
| 154thof 161 | c4ai-command-r-08-2024 | Accuracy | 46.4 |
| 158thof 161 | Llama 3 8B Instruct | Accuracy | 43.93 |
| 159thof 161 | open-mistral-7b | Accuracy | 42.54 |
| 160thof 161 | Llama-2-13b-chat | Accuracy | 42.24 |
| 161stof 161 | Llama-2-70b-chat | Accuracy | 41.65 |
Read from the board on 2026-09-16