Humanity's Last Exam (Epoch AI replication)
Humanity's Last Exam (Epoch AI replication) is run by Epoch AI (AI Benchmarking Hub). It has ranked 51 models, of which the catalogue holds 37, scoring from 2.72 to 46.5 on Accuracy.
measured by Epoch AI (AI Benchmarking Hub) · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 51 | Claude Fable 5.1 | Accuracy (xhigh) | 46.5 |
| 2ndof 51 | Gemini 3.1 Pro Preview | Accuracy | 46.44 |
| 3rdof 51 | GPT-5.4 Pro | Accuracy | 44.32 |
| 4thof 51 | Muse Spark 1.2 | Accuracy | 40.56 |
| 5thof 51 | gemini-3-pro-preview | Accuracy | 37.52 |
| 6thof 51 | GPT-5.4 | Accuracy (xhigh) | 36.24 |
| 7thof 51 | Claude Opus 4.7 | Accuracy | 36.2 |
| 8thof 51 | Claude Opus 4.6 | Accuracy (max) | 34.44 |
| 9thof 51 | GPT-5 Pro | Accuracy | 31.64 |
| 10thof 51 | GPT-5.2 | Accuracy | 27.8 |
| 11thof 51 | GPT-5 | Accuracy (high) | 25.32 |
| 12thof 51 | Claude Opus 4.5 | Accuracy | 25.2 |
| 13thof 51 | Kimi K2.5 | Accuracy | 24.37 |
| 14thof 51 | GPT-5.1 | Accuracy | 23.68 |
| 15thof 51 | Gemini 2.5 Pro Preview 06-05 | Accuracy | 21.64 |
| 16thof 51 | o3 | Accuracy (high) | 20.32 |
| 17thof 51 | GPT-5 mini | Accuracy | 19.44 |
| 21stof 51 | o4-mini | Accuracy (high) | 18.08 |
| 22ndof 51 | Gemini 2.5 Pro Preview 05-06 | Accuracy | 17.8 |
| 25thof 51 | Claude Sonnet 4.5 | Accuracy | 13.72 |
| 26thof 51 | Gemini 2.5 Flash | Accuracy | 12.08 |
| 27thof 51 | Claude Opus 4.1 | Accuracy | 11.52 |
| 29thof 51 | Claude Opus 4 | Accuracy | 10.72 |
| 30thof 51 | Gemini 3.1 Flash-Lite | Accuracy | 8.64 |
| 31stof 51 | GLM-4.5 | Accuracy | 8.32 |
| 32ndof 51 | GLM-4.5-Air | Accuracy | 8.12 |
| 33rdof 51 | o1-pro | Accuracy | 8.12 |
| 34thof 51 | Claude 3.7 Sonnet | Accuracy | 8.04 |
| 35thof 51 | o1 | Accuracy | 7.96 |
| 37thof 51 | Claude Sonnet 4 | Accuracy | 7.76 |
| 42ndof 51 | Llama-4-Maverick-17B-128E-Instruct | Accuracy | 5.68 |
| 45thof 51 | GPT-4.1 fine-tuned | Accuracy | 5.4 |
| 47thof 51 | Mistral Medium 2505 | Accuracy | 4.52 |
| 48thof 51 | Nova Pro | Accuracy | 4.4 |
| 49thof 51 | Claude 3.5 Sonnet | Accuracy | 4.08 |
| 50thof 51 | Nova Lite | Accuracy | 3.64 |
| 51stof 51 | GPT-4o (2024-11-20) | Accuracy | 2.72 |
Read from the board on 2026-09-16