Pass IndexThe State of AISign in

FrontierMath (Tiers 1-3)

FrontierMath (Tiers 1-3) is run by Epoch AI (AI Benchmarking Hub). It has ranked 101 models, of which the catalogue holds 63, scoring from 1.034 to 69 on Accuracy.

measured by Epoch AI (AI Benchmarking Hub) · the board itself

PlaceModelMetricScore
1stof 101GPT-5.5 Promean_score (high)52.4
2ndof 101GPT-5.5mean_score (xhigh)51.7
4thof 101GPT-5.4 Promean_score (xhigh)50
5thof 101GPT-5.4mean_score (xhigh)47.6
6thof 101Claude Opus 4.8mean_score (max)47.241
7thof 101Claude Opus 4.7mean_score (xhigh)43.793
8thof 101Claude Opus 4.6mean_score (max)40.7
9thof 101GPT-5.2mean_score (xhigh)40.7
13thof 101Muse Spark 1.2mean_score39
14thof 101Gemini 3.5 Flashmean_score (high)38.966
15thof 101Kimi K2.6mean_score38.966
17thof 101gemini-3-pro-previewmean_score37.6
18thof 101Gemini 3.1 Pro Previewmean_score36.9
20thof 101Gemini 3 Flash Previewmean_score35.64
21stof 101GLM 5.1mean_score33.448
22ndof 101GPT-5mean_score (high)32.414
23rdof 101Claude Sonnet 4.6mean_score32.4
24thof 101GPT-5.1mean_score (high)31.034
26thof 101GPT-5.4 minimean_score (high)28.28
27thof 101Kimi K2.5mean_score27.9
28thof 101GPT-5 minimean_score (high)27.241
32ndof 101Qwen3.6 Plusmean_score26.207
33rdof 101GPT-5.4 nanomean_score (high)25.86
34thof 101o4-minimean_score (high)24.828
35thof 101Qwen3.6 Max Previewmean_score (max)23.103
36thof 101DeepSeek v3.2mean_score22.1
37thof 101Kimi K2 Thinkingmean_score21.404
38thof 101Qwen 3.5 Plusmean_score21.034
39thof 101Claude Opus 4.5mean_score20.69
45thof 101o3mean_score (high)18.685
48thof 101GLM-5mean_score16.434
49thof 101Claude Sonnet 4.5mean_score15.225
50thof 101Gemini 2.5 Promean_score14.138
52ndof 101o3-minimean_score (high)12.414
54thof 101Qwen3.6-Flashmean_score10.345
55thof 101Gemini 2.5 Pro Preview 06-05mean_score10.345
58thof 101o1mean_score (high)9.31
59thof 101Qwen3 235B A22B Thinking 2507mean_score8.481
60thof 101GPT-5 nanomean_score (high)8.276
63rdof 101Claude Opus 4.1mean_score7.241
64thof 101Qwen3.5-Flashmean_score6.207
65thof 101Claude Haiku 4.5mean_score5.903
67thof 101Grok 3 Minimean_score (high)5.862
68thof 101GPT-4.1 fine-tunedmean_score5.517
69thof 101Gemini 2.5 Flashmean_score4.844
70thof 101Claude Opus 4mean_score4.483
71stof 101GPT-4.1 minimean_score4.483
74thof 101Claude Sonnet 4mean_score4.138
75thof 101Claude 3.7 Sonnetmean_score4.138
76thof 101GLM-4.6mean_score3.819
77thof 101Grok 3mean_score3.793
82ndof 101GLM-4.7mean_score2.439
84thof 101Claude 3.5 Sonnetmean_score2.069
85thof 101Qwen-Plusmean_score1.724
86thof 101gemini-2.0-flash-001mean_score1.724
87thof 101DeepSeek-V3mean_score1.724
88thof 101o1-minimean_score (medium)1.724
90thof 101GPT-4.1 nanomean_score1.034
93rdof 101Llama 4 Maverick 17Bmean_score69
95thof 101Mistral Medium 2505mean_score34.6
96thof 101GPT-4o (2024-08-06)mean_score34.5
97thof 101Claude Haiku 3.5mean_score34.5
99thof 101GPT-4o (2024-11-20)mean_score34.5

Read from the board on 2026-09-16