Critpt — Epoch AI
Critpt — Epoch AI is run by Epoch AI. It has ranked 177 models, of which the catalogue holds 111, scoring from 0 to 85.7.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 177 | GPT-5.6 Sol | Accuracy (max) | 32.3 |
| 2ndof 177 | GPT 6 Astra | Accuracy (max) | 31.714 |
| 4thof 177 | Claude Fable 5.1 | Accuracy (xhigh) | 31.143 |
| 5thof 177 | GPT-5.5 Pro | Accuracy (xhigh) | 30.571 |
| 7thof 177 | GPT-5.4 Pro | Accuracy (xhigh) | 30 |
| 8thof 177 | GPT-5.6 Terra | Accuracy (max) | 30 |
| 12thof 177 | Claude Opus 5 | Accuracy (max) | 29.143 |
| 15thof 177 | Claude Fable 5 | Accuracy (max) | 28.571 |
| 19thof 177 | GPT-5.5 | Accuracy (xhigh) | 27.143 |
| 23rdof 177 | Gemini 3 Deep Think | Accuracy | 25.714 |
| 26thof 177 | GPT-5.4 | Accuracy (xhigh) | 23.429 |
| 27thof 177 | Kimi K3 | Accuracy (max) | 23.4 |
| 31stof 177 | Claude Opus 4.8 | Accuracy (max) | 20.857 |
| 32ndof 177 | GLM 5.2 | Accuracy (max) | 20.857 |
| 33rdof 177 | GPT-5.6 Luna | Accuracy (max) | 20.6 |
| 35thof 177 | Qwen 3.8 Max | Accuracy | 20 |
| 36thof 177 | Grok 4.6 | Accuracy (xhigh) | 19.714 |
| 38thof 177 | GLM-5.3 | Accuracy (max) | 19.143 |
| 40thof 177 | Gemini 3.8 Flash | Accuracy (high) | 18.286 |
| 41stof 177 | DeepSeek V4 Pro 0813 | Accuracy (max) | 18 |
| 42ndof 177 | Gemini 3.1 Pro Preview | Accuracy | 17.714 |
| 44thof 177 | Muse Spark 1.2 | Accuracy (xhigh) | 17.714 |
| 47thof 177 | Claude Sonnet 5 | Accuracy (max) | 16.857 |
| 49thof 177 | DeepSeek V4 Flash (0731) | Accuracy (max) | 16.571 |
| 50thof 177 | Grok 4.5 | Accuracy (high) | 15.429 |
| 51stof 177 | GLM 5.3 Flash | Accuracy | 15.429 |
| 52ndof 177 | Muse Spark 1.1 | Accuracy | 15.1 |
| 54thof 177 | Gemini 3.7 Flash | Accuracy (high) | 14.286 |
| 55thof 177 | Qwen3.7 Max | Accuracy (max) | 13.429 |
| 56thof 177 | Gemini 3.5 Flash | Accuracy (high) | 13.143 |
| 57thof 177 | DeepSeek V4 Pro | Accuracy (max) | 12.857 |
| 58thof 177 | GPT-5 | Accuracy (high) | 12.6 |
| 60thof 177 | Claude Opus 4.7 | Accuracy (max) | 12 |
| 62ndof 177 | Gemini 3.6 Flash | Accuracy (high) | 10.571 |
| 63rdof 177 | GPT-5.4 mini | Accuracy (xhigh) | 10 |
| 64thof 177 | Kimi K2.7 Code | Accuracy | 10 |
| 68thof 177 | GPT-5.4 nano | Accuracy (xhigh) | 9.253 |
| 69thof 177 | Qwen 3.7 Plus | Accuracy | 9.143 |
| 70thof 177 | Grok Build 0.1 | Accuracy | 9.143 |
| 71stof 177 | Inkling Small | Accuracy | 8.286 |
| 73rdof 177 | Kimi K2.6 | Accuracy | 8 |
| 74thof 177 | Grok 4.3 | Accuracy (high) | 8 |
| 75thof 177 | DeepSeek V4 Flash | Accuracy (max) | 7.143 |
| 76thof 177 | gemini-3-pro-preview | Accuracy | 6.9 |
| 79thof 177 | Inkling | Accuracy (xhigh) | 5.429 |
| 80thof 177 | Qwen3.8 27B | Accuracy (xhigh) | 5.429 |
| 82ndof 177 | GPT-5.1 | Accuracy | 4.857 |
| 84thof 177 | GLM 5.1 | Accuracy | 4.571 |
| 85thof 177 | MiMo-V2.5-Pro | Accuracy | 4 |
| 87thof 177 | MiniMax M3 | Accuracy | 3.714 |
| 88thof 177 | Ring-2.6-1T | Accuracy | 3.714 |
| 89thof 177 | MiMo-V2.5 | Accuracy | 3.714 |
| 91stof 177 | Nemotron 3 Ultra | Accuracy | 3.143 |
| 92ndof 177 | Claude Sonnet 4.6 | Accuracy (max) | 3.143 |
| 93rdof 177 | Nemotron 3 Super | Accuracy | 3.143 |
| 94thof 177 | Kimi K2.5 | Accuracy | 3.143 |
| 97thof 177 | Qwen3.6 Plus | Accuracy | 2.857 |
| 100thof 177 | Step 3.7 Flash | Accuracy | 2.286 |
| 101stof 177 | Gemini 2.5 Pro | Accuracy | 2 |
| 103rdof 177 | DeepSeek V3.1 Terminus | Accuracy | 1.714 |
| 104thof 177 | GLM-4.7 | Accuracy | 1.71 |
| 105thof 177 | GPT OSS 20B | Accuracy (high) | 1.429 |
| 107thof 177 | Gemma 4 31B | Accuracy | 1.429 |
| 108thof 177 | o3 | Accuracy (high) | 1.4 |
| 109thof 177 | Gemini 3.1 Flash-Lite | Accuracy | 1.143 |
| 110thof 177 | GPT OSS 120B | Accuracy (high) | 1.143 |
| 111thof 177 | Claude Sonnet 4.5 | Accuracy | 1.143 |
| 112thof 177 | GLM-4.6 | Accuracy | 1.143 |
| 114thof 177 | DeepSeek-R1 | Accuracy | 1.1 |
| 115thof 177 | Gemini 2.5 Flash | Accuracy | 1.1 |
| 116thof 177 | Trinity Large Thinking | Accuracy | 85.7 |
| 117thof 177 | Qwen3.6 27B | Accuracy (none) | 85.7 |
| 118thof 177 | Qwen3.5-122B-A10B | Accuracy (none) | 85.7 |
| 119thof 177 | Mercury 2 | Accuracy | 84.9 |
| 120thof 177 | o4-mini | Accuracy (high) | 60 |
| 121stof 177 | MiniMax M2.7 | Accuracy | 57.1 |
| 122ndof 177 | Qwen3.5-35B-A3B | Accuracy (none) | 57.1 |
| 123rdof 177 | Claude Opus 4 | Accuracy | 30 |
| 124thof 177 | Qwen3.5 9B | Accuracy | 28.6 |
| 125thof 177 | Qwen3.6 35B A3B | Accuracy | 28.6 |
| 126thof 177 | Claude Sonnet 4 | Accuracy | 28.6 |
| 127thof 177 | Qwen3 32B | Accuracy | 28.6 |
| 130thof 177 | Command A Plus | Accuracy | 28.6 |
| 131stof 177 | o3-mini | Accuracy (high) | 28.6 |
| 134thof 177 | Qwen3 30B A3B Thinking 2507 | Accuracy | 28.6 |
| 135thof 177 | Llama-4-Maverick-17B-128E-Instruct | Accuracy | 0 |
| 137thof 177 | GPT-4o (2024-11-20) | Accuracy | 0 |
| 140thof 177 | Claude Haiku 4.5 | Accuracy | 0 |
| 141stof 177 | GPT-4.1 nano | Accuracy | 0 |
| 142ndof 177 | MiMo-V2-Flash | Accuracy | 0 |
| 143rdof 177 | Gemma 4 26B A4B | Accuracy | 0 |
| 144thof 177 | GPT-4.1 mini | Accuracy | 0 |
| 145thof 177 | Gemma 3 27B | Accuracy | 0 |
| 146thof 177 | Llama 3.1 8B Instruct | Accuracy | 0 |
| 147thof 177 | Claude Haiku 3.5 | Accuracy | 0 |
| 148thof 177 | mistral-small-2503 | Accuracy | 0 |
| 149thof 177 | DeepSeek-V3 | Accuracy | 0 |
| 151stof 177 | DeepSeek-V3 0324 | Accuracy | 0 |
| 152ndof 177 | Gemini 3.5 Flash-Lite | Accuracy | 0 |
| 153rdof 177 | Llama 3.3 70B Instruct | Accuracy | 0 |
| 157thof 177 | Mistral Large 3 | Accuracy | 0 |
| 159thof 177 | Solar Pro 3 | Accuracy | 0 |
| 160thof 177 | Phi-4-mini-instruct | Accuracy | 0 |
| 161stof 177 | Qwen3 Coder Next | Accuracy | 0 |
| 162ndof 177 | granite-4.1-30b | Accuracy | 0 |
| 164thof 177 | Gemma 3 12B | Accuracy | 0 |
| 165thof 177 | Qwen3 235B A22B Thinking 2507 | Accuracy | 0 |
| 166thof 177 | GPT-5 mini | Accuracy | 0 |
| 169thof 177 | GPT-5.5 Instant | Accuracy | 0 |
| 175thof 177 | Qwen3 8B | Accuracy | 0 |
| 176thof 177 | Qwen3 14B | Accuracy | 0 |
Read from the board on 2026-09-16