Pass IndexThe State of AISign in

Math level 5 — Epoch AI

Math level 5 — Epoch AI is run by Epoch AI. It has ranked 108 models, of which the catalogue holds 72, scoring from 3.285 to 98.131.

measured by Epoch AI · the board itself

PlaceModelMetricScore
1stof 108GPT-5mean_score (high)98.131
3rdof 108GPT-5 minimean_score (high)97.847
4thof 108o4-minimean_score (high)97.829
5thof 108o3mean_score (high)97.772
6thof 108Claude Sonnet 4.5mean_score97.734
9thof 108DeepSeek-R1 (0528)mean_score96.639
10thof 108o3-minimean_score (high)96.488
11thof 108Claude Haiku 4.5mean_score96.361
12thof 108Gemini 2.5 Pro Preview 05-06mean_score95.903
13thof 108Gemini 2.5 Promean_score95.563
14thof 108GPT-5 nanomean_score (medium)95.242
17thof 108o1mean_score (high)94.713
19thof 108DeepSeek-R1mean_score93.051
20thof 108Claude 3.7 Sonnetmean_score91.163
21stof 108Grok 3 Minimean_score (low)90.937
23rdof 108R1 Distill Llama 70Bmean_score89.898
24thof 108o1-minimean_score (high)89.181
25thof 108Grok 3mean_score88.746
27thof 108GPT-4.1 minimean_score87.292
28thof 108DeepSeek R1 Distill QWEN 14Bmean_score87.122
31stof 108Claude Opus 4mean_score85.045
32ndof 108Claude Sonnet 4mean_score84.366
35thof 108GPT-4.1 fine-tunedmean_score83.006
36thof 108gemini-2.0-flash-001mean_score82.166
38thof 108Mistral Medium 2505mean_score81.628
40thof 108DeepSeek-V3 0324mean_score75.548
41stof 108Gemma 3 27Bmean_score74.037
42ndof 108Llama 4 Maverick 17Bmean_score73.017
44thof 108GPT-4.1 nanomean_score69.996
45thof 108Qwen3 235B A22Bmean_score68.857
48thof 108Qwen-Plusmean_score65.276
49thof 108Phi-4mean_score64.936
50thof 108DeepSeek-V3mean_score64.851
52ndof 108Qwen2.5 72B Instructmean_score63.17
55thof 108Claude 3.5 Sonnetmean_score56.949
56thof 108qwen-turbomean_score56.231
57thof 108AgentRLmean_score56.071
58thof 108GPT-4o (2024-08-06)mean_score53.276
59thof 108GPT-4o-mini (2024-07-18)mean_score52.634
61stof 108GPT-4o (2024-05-13)mean_score51.048
63rdof 108GPT-4o (2024-11-20)mean_score49.773
64thof 108Meta-Llama-3.1-405B-Instructmean_score49.773
65thof 108mistral-small-2503mean_score46.771
66thof 108GPT-4 Turbomean_score46.733
67thof 108Claude Haiku 3.5mean_score46.356
69thof 108Mistral Large 2407mean_score44.817
71stof 108Llama 3.3 70B Instructmean_score41.597
74thof 108Llama-3.2-90B-Vision-Instructmean_score39.435
75thof 108Qwen2-72B-Instructmean_score39.067
76thof 108Claude 3 Opusmean_score37.481
77thof 108Meta-Llama-3.1-70B-Instructmean_score36.679
79thof 108Gemma 2 27Bmean_score27.889
80thof 108WizardLM-2 8x22Bmean_score25.736
81stof 108Yi-1.5-34B-Chatmean_score25.481
83rdof 108Mistral Large (24.02)mean_score24.462
85thof 108GPT-4 (0613)mean_score22.97
86thof 108Llama 3.1 8B Instructmean_score22.876
88thof 108Meta-Llama-3-70B-Instructmean_score22.555
89thof 108gemma-2-9b-itmean_score21.006
90thof 108Claude 3 Sonnetmean_score18.174
91stof 108Phi-3-medium-128k-instructmean_score17.56
92ndof 108GPT-3.5 Turbo (1106)mean_score15.889
94thof 108Claude 3 Haikumean_score14.879
96thof 108Claude 2.0mean_score11.726
98thof 108GPT-3.5 Turbomean_score11.631
102ndof 108Mixtral-8x7B-Instruct-v0.1mean_score9.29
103rdof 108deepseek-llm-67b-chatmean_score6.392
104thof 108Llama 3 8B Instructmean_score6.127
105thof 108Yi-34B-Chatmean_score5.145
106thof 108open-mistral-7bmean_score3.682
107thof 108Mistral-7B-Instruct-v0.3mean_score3.597
108thof 108Llama-2-70b-chat-hfmean_score3.285

Read from the board on 2026-09-16