Pass IndexThe State of AISign in

Humanity's Last Exam (Epoch AI replication)

Humanity's Last Exam (Epoch AI replication) is run by Epoch AI (AI Benchmarking Hub). It has ranked 51 models, of which the catalogue holds 37, scoring from 2.72 to 46.5 on Accuracy.

measured by Epoch AI (AI Benchmarking Hub) · the board itself

PlaceModelMetricScore
1stof 51Claude Fable 5.1Accuracy (xhigh)46.5
2ndof 51Gemini 3.1 Pro PreviewAccuracy46.44
3rdof 51GPT-5.4 ProAccuracy44.32
4thof 51Muse Spark 1.2Accuracy40.56
5thof 51gemini-3-pro-previewAccuracy37.52
6thof 51GPT-5.4Accuracy (xhigh)36.24
7thof 51Claude Opus 4.7Accuracy36.2
8thof 51Claude Opus 4.6Accuracy (max)34.44
9thof 51GPT-5 ProAccuracy31.64
10thof 51GPT-5.2Accuracy27.8
11thof 51GPT-5Accuracy (high)25.32
12thof 51Claude Opus 4.5Accuracy25.2
13thof 51Kimi K2.5Accuracy24.37
14thof 51GPT-5.1Accuracy23.68
15thof 51Gemini 2.5 Pro Preview 06-05Accuracy21.64
16thof 51o3Accuracy (high)20.32
17thof 51GPT-5 miniAccuracy19.44
21stof 51o4-miniAccuracy (high)18.08
22ndof 51Gemini 2.5 Pro Preview 05-06Accuracy17.8
25thof 51Claude Sonnet 4.5Accuracy13.72
26thof 51Gemini 2.5 FlashAccuracy12.08
27thof 51Claude Opus 4.1Accuracy11.52
29thof 51Claude Opus 4Accuracy10.72
30thof 51Gemini 3.1 Flash-LiteAccuracy8.64
31stof 51GLM-4.5Accuracy8.32
32ndof 51GLM-4.5-AirAccuracy8.12
33rdof 51o1-proAccuracy8.12
34thof 51Claude 3.7 SonnetAccuracy8.04
35thof 51o1Accuracy7.96
37thof 51Claude Sonnet 4Accuracy7.76
42ndof 51Llama-4-Maverick-17B-128E-InstructAccuracy5.68
45thof 51GPT-4.1 fine-tunedAccuracy5.4
47thof 51Mistral Medium 2505Accuracy4.52
48thof 51Nova ProAccuracy4.4
49thof 51Claude 3.5 SonnetAccuracy4.08
50thof 51Nova LiteAccuracy3.64
51stof 51GPT-4o (2024-11-20)Accuracy2.72

Read from the board on 2026-09-16