Weirdml — Epoch AI
Weirdml — Epoch AI is run by Epoch AI. It has ranked 170 models, of which the catalogue holds 117, scoring from 1.73 to 92.9.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 170 | Claude Fable 5.1 | Accuracy (max) | 92.9 |
| 2ndof 170 | GPT 6 Astra | Accuracy (high) | 92.87 |
| 4thof 170 | Claude Fable 5 | Accuracy (max) | 91.94 |
| 5thof 170 | Claude Opus 5 | Accuracy (max) | 91.78 |
| 8thof 170 | GPT-5.6 Sol | Accuracy (high) | 88.76 |
| 12thof 170 | GPT-5.5 | Accuracy (xhigh) | 84.91 |
| 14thof 170 | Claude Opus 4.8 | Accuracy (xhigh) | 82.89 |
| 15thof 170 | Kimi K3 | Accuracy (max) | 82.57 |
| 16thof 170 | GPT-5.3 Codex | Accuracy | 79.3 |
| 17thof 170 | GPT-5.6 Terra | Accuracy (high) | 78.27 |
| 18thof 170 | Claude Opus 4.6 | Accuracy (high) | 77.95 |
| 21stof 170 | GPT-5.4 | Accuracy (xhigh) | 77.7 |
| 22ndof 170 | Claude Opus 4.7 | Accuracy (high) | 76.44 |
| 27thof 170 | GLM-5.3 | Accuracy (max) | 75.4 |
| 28thof 170 | GPT-5.2 | Accuracy (xhigh) | 72.19 |
| 29thof 170 | Gemini 3.1 Pro Preview | Accuracy | 72.07 |
| 31stof 170 | GLM 5.2 | Accuracy (max) | 70.12 |
| 32ndof 170 | gemini-3-pro-preview | Accuracy | 69.93 |
| 33rdof 170 | Claude Sonnet 5 | Accuracy (high) | 68.78 |
| 35thof 170 | Grok 4.6 | Accuracy (high) | 67.29 |
| 37thof 170 | DeepSeek V4 Pro 0813 | Accuracy (max) | 66.22 |
| 38thof 170 | Claude Sonnet 4.6 | Accuracy (medium) | 66.07 |
| 40thof 170 | Claude Opus 4.5 | Accuracy | 63.74 |
| 42ndof 170 | DeepSeek V4 Flash (0731) | Accuracy (max) | 62.96 |
| 43rdof 170 | Gemini 3.5 Flash | Accuracy (high) | 62.64 |
| 44thof 170 | Gemini 3 Flash Preview | Accuracy | 61.6 |
| 45thof 170 | GPT-5.6 Luna | Accuracy (high) | 60.86 |
| 46thof 170 | GPT-5.1 | Accuracy (high) | 60.77 |
| 47thof 170 | GPT-5 | Accuracy (high) | 60.7 |
| 48thof 170 | GPT-5 Pro | Accuracy (high) | 60.39 |
| 49thof 170 | GPT-5.4 mini | Accuracy (high) | 60.3 |
| 50thof 170 | Muse Spark 1.2 | Accuracy (xhigh) | 60.3 |
| 51stof 170 | o3-pro | Accuracy (high) | 58.21 |
| 53rdof 170 | GPT-5.4 Pro | Accuracy (none) | 57.44 |
| 54thof 170 | GLM 5.1 | Accuracy | 57.1 |
| 56thof 170 | Gemini 3.6 Flash | Accuracy (high) | 56.1 |
| 57thof 170 | Kimi K2.6 | Accuracy | 55.86 |
| 58thof 170 | GPT-5 Codex (batch) | Accuracy | 54.53 |
| 60thof 170 | Kimi K2.7 Code | Accuracy | 54.12 |
| 61stof 170 | Gemini 2.5 Pro | Accuracy | 54.03 |
| 62ndof 170 | GPT-5 mini | Accuracy (high) | 52.67 |
| 63rdof 170 | o4-mini | Accuracy (high) | 52.56 |
| 64thof 170 | o3 | Accuracy (high) | 52.42 |
| 65thof 170 | Grok 4.20 | Accuracy | 52.26 |
| 66thof 170 | Gemma 4 31B | Accuracy | 52.26 |
| 67thof 170 | Gemini 3.1 Flash-Lite | Accuracy | 52.19 |
| 68thof 170 | Grok 4.3 | Accuracy | 49.89 |
| 71stof 170 | GPT-5.4 nano | Accuracy (high) | 49.23 |
| 72ndof 170 | DeepSeek V4 Pro | Accuracy (max) | 48.9 |
| 73rdof 170 | GPT OSS 120B | Accuracy (high) | 48.17 |
| 75thof 170 | GLM-5 | Accuracy | 48.17 |
| 76thof 170 | Claude Sonnet 4.5 | Accuracy | 47.71 |
| 77thof 170 | o1 | Accuracy | 47.56 |
| 78thof 170 | DeepSeek-V3.2-Speciale | Accuracy | 46.73 |
| 81stof 170 | Grok 4.5 | Accuracy | 46.43 |
| 82ndof 170 | Claude Sonnet 4 | Accuracy | 46.11 |
| 84thof 170 | Claude Opus 4.1 | Accuracy | 45.86 |
| 86thof 170 | DeepSeek V4 Flash | Accuracy (max) | 45.63 |
| 87thof 170 | Kimi K2.5 | Accuracy | 45.6 |
| 88thof 170 | Claude Haiku 4.5 | Accuracy | 45.4 |
| 92ndof 170 | Claude Opus 4 | Accuracy | 43.72 |
| 94thof 170 | o3-mini | Accuracy (high) | 43.7 |
| 95thof 170 | Nemotron 3 Ultra | Accuracy | 43.45 |
| 96thof 170 | Mercury 2 | Accuracy | 43.2 |
| 97thof 170 | Grok 4 Fast | Accuracy | 42.86 |
| 98thof 170 | Kimi K2 Thinking | Accuracy | 42.79 |
| 99thof 170 | Grok 3 Mini | Accuracy (high) | 42.58 |
| 103rdof 170 | DeepSeek-R1 (0528) | Accuracy | 41.63 |
| 104thof 170 | Qwen3-Coder-480B-A35B-Instruct | Accuracy | 41.17 |
| 105thof 170 | Qwen3 235B A22B Thinking 2507 | Accuracy | 41.04 |
| 106thof 170 | Gemini 2.5 Flash | Accuracy | 40.95 |
| 108thof 170 | GPT OSS 20B | Accuracy (high) | 40.93 |
| 110thof 170 | Claude 3.5 Sonnet | Accuracy | 39.97 |
| 111thof 170 | GPT-5 Chat | Accuracy | 39.77 |
| 112thof 170 | Qwen3.5-27B | Accuracy | 39.53 |
| 115thof 170 | Kimi K2 Instruct | Accuracy | 39.36 |
| 116thof 170 | Kimi K2 0711 | Accuracy | 39.36 |
| 117thof 170 | GPT-4.1 fine-tuned | Accuracy | 39.04 |
| 118thof 170 | Gemini 3.5 Flash-Lite | Accuracy (high) | 39 |
| 119thof 170 | Qwen3-235B-A22B-Instruct-2507 | Accuracy | 38.7 |
| 122ndof 170 | GPT-5 nano | Accuracy (high) | 38.06 |
| 123rdof 170 | Nemotron 3 Super | Accuracy | 38.01 |
| 126thof 170 | GPT-4.1 mini | Accuracy | 37.61 |
| 127thof 170 | DeepSeek-V3.1 | Accuracy | 37.5 |
| 129thof 170 | Qwen3 235B A22B | Accuracy | 37.28 |
| 130thof 170 | Grok 3 | Accuracy | 37.24 |
| 131stof 170 | MiniMax M2.7 | Accuracy | 36.95 |
| 134thof 170 | DeepSeek-R1 | Accuracy | 36.49 |
| 135thof 170 | o1-mini | Accuracy (medium) | 36.32 |
| 136thof 170 | DeepSeek-V3 0324 | Accuracy | 36.08 |
| 137thof 170 | Gemini 2.5 Flash-Lite | Accuracy | 35.22 |
| 139thof 170 | Gemma 4 26B A4B | Accuracy | 35.17 |
| 140thof 170 | grok-code-fast-1 | Accuracy | 35.06 |
| 141stof 170 | Qwen3.6 35B A3B | Accuracy | 34.49 |
| 142ndof 170 | Qwen3 Coder Next | Accuracy | 34.4 |
| 144thof 170 | Inkling | Accuracy (high) | 32.29 |
| 146thof 170 | Claude Haiku 3.5 | Accuracy | 30.73 |
| 147thof 170 | Qwen3 30B A3B | Accuracy | 29.75 |
| 149thof 170 | gemini-2.0-flash-001 | Accuracy | 25.77 |
| 150thof 170 | GPT-4o (2024-11-20) | Accuracy | 25.12 |
| 152ndof 170 | Llama-4-Maverick-17B-128E-Instruct | Accuracy | 24.47 |
| 155thof 170 | Meta-Llama-3.1-405B-Instruct | Accuracy | 21.38 |
| 156thof 170 | Claude 3 Opus | Accuracy | 19.22 |
| 157thof 170 | GPT-4.1 nano | Accuracy | 18.98 |
| 158thof 170 | GPT-4 Turbo | Accuracy | 18.01 |
| 159thof 170 | Qwen2.5 72B Instruct | Accuracy | 15.97 |
| 160thof 170 | Llama 3.3 70B Instruct | Accuracy | 14.44 |
| 161stof 170 | GPT-4 (0613) | Accuracy | 12.36 |
| 162ndof 170 | GPT-4o-mini (2024-07-18) | Accuracy | 11.76 |
| 163rdof 170 | Qwen2-72B-Instruct | Accuracy | 11.3 |
| 164thof 170 | Claude 3 Sonnet | Accuracy | 10.16 |
| 165thof 170 | Claude 3 Haiku | Accuracy | 9.84 |
| 166thof 170 | Meta-Llama-3.1-70B-Instruct | Accuracy | 8.97 |
| 167thof 170 | Claude 2.1 | Accuracy | 7.06 |
| 168thof 170 | GPT-3.5 Turbo | Accuracy | 3.48 |
| 169thof 170 | Mixtral-8x22B-Instruct-v0.1 | Accuracy | 3.17 |
| 170thof 170 | Llama 3.1 8B Instruct | Accuracy | 1.73 |
Read from the board on 2026-09-16