Pass IndexThe State of AISign in

Humanity's Last Exam

Humanity's Last Exam is run by Center for AI Safety + Scale AI (agi.safe.ai / lastexam.ai). It has ranked 10 models, of which the catalogue holds 10, scoring from 2.7 to 38.3 on Accuracy (%).

measured by Center for AI Safety + Scale AI (agi.safe.ai / lastexam.ai) · the board itself

PlaceModelMetricScore
1stof 10gemini-3-pro-previewAccuracy (%)38.3
2ndof 10GPT-5Accuracy (%)25.3
3rdof 10Grok 4Accuracy (%)24.5
4thof 10Gemini 2.5 ProAccuracy (%)21.6
5thof 10GPT-5 miniAccuracy (%)19.4
6thof 10Claude Sonnet 4.5Accuracy (%)13.7
7thof 10Gemini 2.5 FlashAccuracy (%)12.1
8thof 10DeepSeek-R1Accuracy (%)8.5
9thof 10o1Accuracy (%)8
10thof 10GPT-4oAccuracy (%)2.7

Read from the board on 2026-08-25