Math level 5 — Epoch AI
Math level 5 — Epoch AI is run by Epoch AI. It has ranked 108 models, of which the catalogue holds 72, scoring from 3.285 to 98.131.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 108 | GPT-5 | mean_score (high) | 98.131 |
| 3rdof 108 | GPT-5 mini | mean_score (high) | 97.847 |
| 4thof 108 | o4-mini | mean_score (high) | 97.829 |
| 5thof 108 | o3 | mean_score (high) | 97.772 |
| 6thof 108 | Claude Sonnet 4.5 | mean_score | 97.734 |
| 9thof 108 | DeepSeek-R1 (0528) | mean_score | 96.639 |
| 10thof 108 | o3-mini | mean_score (high) | 96.488 |
| 11thof 108 | Claude Haiku 4.5 | mean_score | 96.361 |
| 12thof 108 | Gemini 2.5 Pro Preview 05-06 | mean_score | 95.903 |
| 13thof 108 | Gemini 2.5 Pro | mean_score | 95.563 |
| 14thof 108 | GPT-5 nano | mean_score (medium) | 95.242 |
| 17thof 108 | o1 | mean_score (high) | 94.713 |
| 19thof 108 | DeepSeek-R1 | mean_score | 93.051 |
| 20thof 108 | Claude 3.7 Sonnet | mean_score | 91.163 |
| 21stof 108 | Grok 3 Mini | mean_score (low) | 90.937 |
| 23rdof 108 | R1 Distill Llama 70B | mean_score | 89.898 |
| 24thof 108 | o1-mini | mean_score (high) | 89.181 |
| 25thof 108 | Grok 3 | mean_score | 88.746 |
| 27thof 108 | GPT-4.1 mini | mean_score | 87.292 |
| 28thof 108 | DeepSeek R1 Distill QWEN 14B | mean_score | 87.122 |
| 31stof 108 | Claude Opus 4 | mean_score | 85.045 |
| 32ndof 108 | Claude Sonnet 4 | mean_score | 84.366 |
| 35thof 108 | GPT-4.1 fine-tuned | mean_score | 83.006 |
| 36thof 108 | gemini-2.0-flash-001 | mean_score | 82.166 |
| 38thof 108 | Mistral Medium 2505 | mean_score | 81.628 |
| 40thof 108 | DeepSeek-V3 0324 | mean_score | 75.548 |
| 41stof 108 | Gemma 3 27B | mean_score | 74.037 |
| 42ndof 108 | Llama 4 Maverick 17B | mean_score | 73.017 |
| 44thof 108 | GPT-4.1 nano | mean_score | 69.996 |
| 45thof 108 | Qwen3 235B A22B | mean_score | 68.857 |
| 48thof 108 | Qwen-Plus | mean_score | 65.276 |
| 49thof 108 | Phi-4 | mean_score | 64.936 |
| 50thof 108 | DeepSeek-V3 | mean_score | 64.851 |
| 52ndof 108 | Qwen2.5 72B Instruct | mean_score | 63.17 |
| 55thof 108 | Claude 3.5 Sonnet | mean_score | 56.949 |
| 56thof 108 | qwen-turbo | mean_score | 56.231 |
| 57thof 108 | AgentRL | mean_score | 56.071 |
| 58thof 108 | GPT-4o (2024-08-06) | mean_score | 53.276 |
| 59thof 108 | GPT-4o-mini (2024-07-18) | mean_score | 52.634 |
| 61stof 108 | GPT-4o (2024-05-13) | mean_score | 51.048 |
| 63rdof 108 | GPT-4o (2024-11-20) | mean_score | 49.773 |
| 64thof 108 | Meta-Llama-3.1-405B-Instruct | mean_score | 49.773 |
| 65thof 108 | mistral-small-2503 | mean_score | 46.771 |
| 66thof 108 | GPT-4 Turbo | mean_score | 46.733 |
| 67thof 108 | Claude Haiku 3.5 | mean_score | 46.356 |
| 69thof 108 | Mistral Large 2407 | mean_score | 44.817 |
| 71stof 108 | Llama 3.3 70B Instruct | mean_score | 41.597 |
| 74thof 108 | Llama-3.2-90B-Vision-Instruct | mean_score | 39.435 |
| 75thof 108 | Qwen2-72B-Instruct | mean_score | 39.067 |
| 76thof 108 | Claude 3 Opus | mean_score | 37.481 |
| 77thof 108 | Meta-Llama-3.1-70B-Instruct | mean_score | 36.679 |
| 79thof 108 | Gemma 2 27B | mean_score | 27.889 |
| 80thof 108 | WizardLM-2 8x22B | mean_score | 25.736 |
| 81stof 108 | Yi-1.5-34B-Chat | mean_score | 25.481 |
| 83rdof 108 | Mistral Large (24.02) | mean_score | 24.462 |
| 85thof 108 | GPT-4 (0613) | mean_score | 22.97 |
| 86thof 108 | Llama 3.1 8B Instruct | mean_score | 22.876 |
| 88thof 108 | Meta-Llama-3-70B-Instruct | mean_score | 22.555 |
| 89thof 108 | gemma-2-9b-it | mean_score | 21.006 |
| 90thof 108 | Claude 3 Sonnet | mean_score | 18.174 |
| 91stof 108 | Phi-3-medium-128k-instruct | mean_score | 17.56 |
| 92ndof 108 | GPT-3.5 Turbo (1106) | mean_score | 15.889 |
| 94thof 108 | Claude 3 Haiku | mean_score | 14.879 |
| 96thof 108 | Claude 2.0 | mean_score | 11.726 |
| 98thof 108 | GPT-3.5 Turbo | mean_score | 11.631 |
| 102ndof 108 | Mixtral-8x7B-Instruct-v0.1 | mean_score | 9.29 |
| 103rdof 108 | deepseek-llm-67b-chat | mean_score | 6.392 |
| 104thof 108 | Llama 3 8B Instruct | mean_score | 6.127 |
| 105thof 108 | Yi-34B-Chat | mean_score | 5.145 |
| 106thof 108 | open-mistral-7b | mean_score | 3.682 |
| 107thof 108 | Mistral-7B-Instruct-v0.3 | mean_score | 3.597 |
| 108thof 108 | Llama-2-70b-chat-hf | mean_score | 3.285 |
Read from the board on 2026-09-16