AIME 2025
AIME 2025 is run by MathArena (ETH Zurich SRI Lab). It has ranked 61 models, of which the catalogue holds 23, scoring from 88.33 to 100 on Accuracy (± 95% CI).
measured by MathArena (ETH Zurich SRI Lab) · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 61 | GPT-5.2 | Accuracy (± 95% CI) (xhigh) | 100 |
| 3rdof 61 | Step 3.5 Flash | Accuracy (± 95% CI) | 98.33 |
| 4thof 61 | Gemini 3 Flash Preview | Accuracy (± 95% CI) | 97.5 |
| 5thof 61 | GLM-5 | Accuracy (± 95% CI) | 96.67 |
| 6thof 61 | Kimi K2.5 | Accuracy (± 95% CI) (think) | 95.83 |
| 6thof 61 | DeepSeek-V3.2-Speciale | Accuracy (± 95% CI) | 95.83 |
| 8thof 61 | GPT-5 | Accuracy (± 95% CI) (high) | 95 |
| 8thof 61 | gemini-3-pro-preview | Accuracy (± 95% CI) (preview) | 95 |
| 10thof 61 | DeepSeek v3.2 | Accuracy (± 95% CI) (think) | 94.17 |
| 10thof 61 | GPT-5.1 | Accuracy (± 95% CI) (high) | 94.17 |
| 12thof 61 | GLM-4.5 | Accuracy (± 95% CI) | 93.33 |
| 13thof 61 | Kimi K2 Thinking | Accuracy (± 95% CI) | 92.5 |
| 13thof 61 | Grok 4 | Accuracy (± 95% CI) | 92.5 |
| 15thof 61 | DeepSeek V3.2 Exp | Accuracy (± 95% CI) (think) | 91.67 |
| 15thof 61 | GLM-4.6 | Accuracy (± 95% CI) | 91.67 |
| 15thof 61 | o4 Mini High | Accuracy (± 95% CI) | 91.67 |
| 18thof 61 | DeepSeek-V3.1 | Accuracy (± 95% CI) (think) | 90.83 |
| 20thof 61 | GPT OSS 120B | Accuracy (± 95% CI) (high) | 90 |
| 21stof 61 | Grok 4.1 Fast Reasoning | Accuracy (± 95% CI) | 89.17 |
| 21stof 61 | GPT OSS 20B | Accuracy (± 95% CI) (high) | 89.17 |
| 21stof 61 | DeepSeek-R1 (0528) | Accuracy (± 95% CI) | 89.17 |
| 21stof 61 | o3 | Accuracy (± 95% CI) (high) | 89.17 |
| 25thof 61 | Gemini 2.5 Pro | Accuracy (± 95% CI) | 88.33 |
Read from the board on 2026-08-25