Proofbench — Epoch AI
Proofbench — Epoch AI is run by Epoch AI. It has ranked 64 models, of which the catalogue holds 62, scoring from 0 to 100.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 64 | Claude Fable 5.1 | Accuracy (max) | 100 |
| 2ndof 64 | Claude Opus 5 | Accuracy (max) | 99 |
| 3rdof 64 | GPT 6 Astra | Accuracy | 99 |
| 4thof 64 | Claude Fable 5 | Accuracy (max) | 95 |
| 5thof 64 | Kimi K3 | Accuracy | 87 |
| 6thof 64 | GPT-5.6 Sol | Accuracy (max) | 83 |
| 7thof 64 | Claude Sonnet 5 | Accuracy (max) | 77 |
| 8thof 64 | GPT-5.6 Terra | Accuracy (xhigh) | 74 |
| 9thof 64 | Claude Opus 4.8 | Accuracy (max) | 69 |
| 10thof 64 | GPT-5.6 Luna | Accuracy (max) | 60 |
| 11thof 64 | Gemini 3.7 Flash | Accuracy | 58 |
| 12thof 64 | Qwen 3.8 Max | Accuracy | 58 |
| 13thof 64 | GPT-5.4 | Accuracy (xhigh) | 56 |
| 14thof 64 | DeepSeek V4 Flash (0731) | Accuracy | 56 |
| 15thof 64 | Claude Opus 4.7 | Accuracy (max) | 54 |
| 16thof 64 | Grok 4.6 | Accuracy | 51 |
| 17thof 64 | GPT-5.5 | Accuracy (xhigh) | 50 |
| 18thof 64 | Claude Opus 4.6 | Accuracy (max) | 50 |
| 19thof 64 | DeepSeek V4 Pro 0813 | Accuracy | 50 |
| 20thof 64 | GLM-5.3 | Accuracy (max) | 49 |
| 21stof 64 | Gemini 3.8 Flash | Accuracy | 48 |
| 22ndof 64 | Claude Sonnet 4.6 | Accuracy (max) | 45 |
| 23rdof 64 | Muse Spark 1.2 | Accuracy | 43 |
| 24thof 64 | Muse Spark 1.1 | Accuracy | 39 |
| 25thof 64 | Claude Opus 4.5 | Accuracy | 36 |
| 26thof 64 | Gemini 3.6 Flash | Accuracy | 36 |
| 27thof 64 | GLM 5.2 | Accuracy (max) | 35 |
| 28thof 64 | Grok 4.5 | Accuracy (high) | 31 |
| 29thof 64 | Gemini 3.5 Flash | Accuracy (high) | 31 |
| 30thof 64 | Qwen3.7 Max | Accuracy (max) | 26 |
| 31stof 64 | Gemini 3.1 Pro Preview | Accuracy | 26 |
| 32ndof 64 | GLM 5.1 | Accuracy | 22.222 |
| 33rdof 64 | MiMo-V2.5-Pro | Accuracy | 22 |
| 34thof 64 | GPT-5.4 mini | Accuracy (xhigh) | 21 |
| 35thof 64 | GLM 5.3 Flash | Accuracy (max) | 21 |
| 36thof 64 | gemini-3-pro-preview | Accuracy | 20 |
| 37thof 64 | Claude Sonnet 4.5 | Accuracy | 19 |
| 38thof 64 | MiniMax M3 | Accuracy | 18 |
| 39thof 64 | GPT-5 | Accuracy (high) | 18 |
| 41stof 64 | Kimi K2.6 | Accuracy | 16 |
| 42ndof 64 | MiMo-V2.5 | Accuracy | 16 |
| 43rdof 64 | DeepSeek V4 Pro | Accuracy (max) | 16 |
| 44thof 64 | Qwen3.8 27B | Accuracy | 16 |
| 45thof 64 | Gemini 3 Flash Preview | Accuracy | 15 |
| 46thof 64 | GPT-5.2 | Accuracy (xhigh) | 15 |
| 47thof 64 | Grok 4.20 | Accuracy | 14 |
| 48thof 64 | Gemini 3.5 Flash-Lite | Accuracy | 13 |
| 49thof 64 | GPT-5 nano | Accuracy (high) | 12 |
| 50thof 64 | Grok 4.3 | Accuracy (high) | 11 |
| 51stof 64 | GPT-5 mini | Accuracy (high) | 9 |
| 52ndof 64 | GPT-5.1 Codex Max | Accuracy (max) | 9 |
| 54thof 64 | DeepSeek v3.2 | Accuracy | 8 |
| 55thof 64 | GLM-4.7 | Accuracy | 6 |
| 56thof 64 | Inkling Small | Accuracy | 6 |
| 57thof 64 | GPT-5.4 nano | Accuracy (high) | 5 |
| 58thof 64 | Grok 4.1 Fast Reasoning | Accuracy | 4 |
| 59thof 64 | MiniMax M2.5 | Accuracy | 4 |
| 60thof 64 | MiniMax M2.7 | Accuracy | 3 |
| 61stof 64 | Nemotron 3 Ultra | Accuracy | 2 |
| 62ndof 64 | Laguna-XS.2 | Accuracy | 0 |
| 63rdof 64 | Laguna M.1 | Accuracy | 0 |
| 64thof 64 | Inkling | Accuracy | 0 |
Read from the board on 2026-09-16