FrontierMath (Tiers 1-3)
FrontierMath (Tiers 1-3) is run by Epoch AI (AI Benchmarking Hub). It has ranked 101 models, of which the catalogue holds 63, scoring from 1.034 to 69 on Accuracy.
measured by Epoch AI (AI Benchmarking Hub) · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 101 | GPT-5.5 Pro | mean_score (high) | 52.4 |
| 2ndof 101 | GPT-5.5 | mean_score (xhigh) | 51.7 |
| 4thof 101 | GPT-5.4 Pro | mean_score (xhigh) | 50 |
| 5thof 101 | GPT-5.4 | mean_score (xhigh) | 47.6 |
| 6thof 101 | Claude Opus 4.8 | mean_score (max) | 47.241 |
| 7thof 101 | Claude Opus 4.7 | mean_score (xhigh) | 43.793 |
| 8thof 101 | Claude Opus 4.6 | mean_score (max) | 40.7 |
| 9thof 101 | GPT-5.2 | mean_score (xhigh) | 40.7 |
| 13thof 101 | Muse Spark 1.2 | mean_score | 39 |
| 14thof 101 | Gemini 3.5 Flash | mean_score (high) | 38.966 |
| 15thof 101 | Kimi K2.6 | mean_score | 38.966 |
| 17thof 101 | gemini-3-pro-preview | mean_score | 37.6 |
| 18thof 101 | Gemini 3.1 Pro Preview | mean_score | 36.9 |
| 20thof 101 | Gemini 3 Flash Preview | mean_score | 35.64 |
| 21stof 101 | GLM 5.1 | mean_score | 33.448 |
| 22ndof 101 | GPT-5 | mean_score (high) | 32.414 |
| 23rdof 101 | Claude Sonnet 4.6 | mean_score | 32.4 |
| 24thof 101 | GPT-5.1 | mean_score (high) | 31.034 |
| 26thof 101 | GPT-5.4 mini | mean_score (high) | 28.28 |
| 27thof 101 | Kimi K2.5 | mean_score | 27.9 |
| 28thof 101 | GPT-5 mini | mean_score (high) | 27.241 |
| 32ndof 101 | Qwen3.6 Plus | mean_score | 26.207 |
| 33rdof 101 | GPT-5.4 nano | mean_score (high) | 25.86 |
| 34thof 101 | o4-mini | mean_score (high) | 24.828 |
| 35thof 101 | Qwen3.6 Max Preview | mean_score (max) | 23.103 |
| 36thof 101 | DeepSeek v3.2 | mean_score | 22.1 |
| 37thof 101 | Kimi K2 Thinking | mean_score | 21.404 |
| 38thof 101 | Qwen 3.5 Plus | mean_score | 21.034 |
| 39thof 101 | Claude Opus 4.5 | mean_score | 20.69 |
| 45thof 101 | o3 | mean_score (high) | 18.685 |
| 48thof 101 | GLM-5 | mean_score | 16.434 |
| 49thof 101 | Claude Sonnet 4.5 | mean_score | 15.225 |
| 50thof 101 | Gemini 2.5 Pro | mean_score | 14.138 |
| 52ndof 101 | o3-mini | mean_score (high) | 12.414 |
| 54thof 101 | Qwen3.6-Flash | mean_score | 10.345 |
| 55thof 101 | Gemini 2.5 Pro Preview 06-05 | mean_score | 10.345 |
| 58thof 101 | o1 | mean_score (high) | 9.31 |
| 59thof 101 | Qwen3 235B A22B Thinking 2507 | mean_score | 8.481 |
| 60thof 101 | GPT-5 nano | mean_score (high) | 8.276 |
| 63rdof 101 | Claude Opus 4.1 | mean_score | 7.241 |
| 64thof 101 | Qwen3.5-Flash | mean_score | 6.207 |
| 65thof 101 | Claude Haiku 4.5 | mean_score | 5.903 |
| 67thof 101 | Grok 3 Mini | mean_score (high) | 5.862 |
| 68thof 101 | GPT-4.1 fine-tuned | mean_score | 5.517 |
| 69thof 101 | Gemini 2.5 Flash | mean_score | 4.844 |
| 70thof 101 | Claude Opus 4 | mean_score | 4.483 |
| 71stof 101 | GPT-4.1 mini | mean_score | 4.483 |
| 74thof 101 | Claude Sonnet 4 | mean_score | 4.138 |
| 75thof 101 | Claude 3.7 Sonnet | mean_score | 4.138 |
| 76thof 101 | GLM-4.6 | mean_score | 3.819 |
| 77thof 101 | Grok 3 | mean_score | 3.793 |
| 82ndof 101 | GLM-4.7 | mean_score | 2.439 |
| 84thof 101 | Claude 3.5 Sonnet | mean_score | 2.069 |
| 85thof 101 | Qwen-Plus | mean_score | 1.724 |
| 86thof 101 | gemini-2.0-flash-001 | mean_score | 1.724 |
| 87thof 101 | DeepSeek-V3 | mean_score | 1.724 |
| 88thof 101 | o1-mini | mean_score (medium) | 1.724 |
| 90thof 101 | GPT-4.1 nano | mean_score | 1.034 |
| 93rdof 101 | Llama 4 Maverick 17B | mean_score | 69 |
| 95thof 101 | Mistral Medium 2505 | mean_score | 34.6 |
| 96thof 101 | GPT-4o (2024-08-06) | mean_score | 34.5 |
| 97thof 101 | Claude Haiku 3.5 | mean_score | 34.5 |
| 99thof 101 | GPT-4o (2024-11-20) | mean_score | 34.5 |
Read from the board on 2026-09-16