Humanity's Last Exam
Humanity's Last Exam is run by Center for AI Safety + Scale AI (agi.safe.ai / lastexam.ai). It has ranked 10 models, of which the catalogue holds 10, scoring from 2.7 to 38.3 on Accuracy (%).
measured by Center for AI Safety + Scale AI (agi.safe.ai / lastexam.ai) · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 10 | gemini-3-pro-preview | Accuracy (%) | 38.3 |
| 2ndof 10 | GPT-5 | Accuracy (%) | 25.3 |
| 3rdof 10 | Grok 4 | Accuracy (%) | 24.5 |
| 4thof 10 | Gemini 2.5 Pro | Accuracy (%) | 21.6 |
| 5thof 10 | GPT-5 mini | Accuracy (%) | 19.4 |
| 6thof 10 | Claude Sonnet 4.5 | Accuracy (%) | 13.7 |
| 7thof 10 | Gemini 2.5 Flash | Accuracy (%) | 12.1 |
| 8thof 10 | DeepSeek-R1 | Accuracy (%) | 8.5 |
| 9thof 10 | o1 | Accuracy (%) | 8 |
| 10thof 10 | GPT-4o | Accuracy (%) | 2.7 |
Read from the board on 2026-08-25