MMLU-Pro
MMLU-Pro is run by TIGER-Lab. It has ranked 262 models, of which the catalogue holds 21, scoring from 86.4 to 91.16 on Overall (accuracy).
measured by TIGER-Lab · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 262 | Gemini 3.1 Pro Preview | Overall (accuracy) | 91.16 |
| 2ndof 262 | gemini-3-pro-preview | Overall (accuracy) (11/25) | 90.1 |
| 3rdof 262 | o1 | Overall (accuracy) | 89.3 |
| 4thof 262 | Claude Opus 4.6 | Overall (accuracy) (thinking) | 89.1 |
| 5thof 262 | Gemini 3 Flash Preview | Overall (accuracy) (12/25) | 88.6 |
| 6thof 262 | MiniMax M2.1 | Overall (accuracy) | 88 |
| 7thof 262 | Qwen3.5 397B A17B | Overall (accuracy) | 87.8 |
| 8thof 262 | Seed-2.0-Lite | Overall (accuracy) | 87.7 |
| 9thof 262 | GPT-5.4 | Overall (accuracy) | 87.5 |
| 10thof 262 | GPT-5.2 | Overall (accuracy) | 87.4 |
| 11thof 262 | Claude Sonnet 4.5 | Overall (accuracy) (thinking) | 87.4 |
| 12thof 262 | Claude Opus 4 | Overall (accuracy) (thinking) | 87.3 |
| 13thof 262 | Claude Opus 4.5 | Overall (accuracy) (thinking) | 87.3 |
| 14thof 262 | Claude Sonnet 4.6 | Overall (accuracy) (thinking) | 87.3 |
| 16thof 262 | GPT-5 | Overall (accuracy) (high) | 87.1 |
| 17thof 262 | Kimi K2.5 | Overall (accuracy) | 87.1 |
| 19thof 262 | Grok 4 | Overall (accuracy) | 87 |
| 20thof 262 | Seed-2.0-pro | Overall (accuracy) | 87 |
| 21stof 262 | Qwen3.5-122B-A10B | Overall (accuracy) | 86.7 |
| 22ndof 262 | Seed 1.6 | Overall (accuracy) (thinking) | 86.6 |
| 25thof 262 | GPT-5.1 | Overall (accuracy) | 86.4 |
Read from the board on 2026-08-25