Frontiermath tier 4 — Epoch AI
Frontiermath tier 4 — Epoch AI is run by Epoch AI. It has ranked 72 models, of which the catalogue holds 49, scoring from 0 to 39.6.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 2ndof 72 | GPT-5.5 Pro | mean_score (xhigh) | 39.6 |
| 5thof 72 | GPT-5.5 | mean_score (xhigh) | 35.4 |
| 7thof 72 | Claude Opus 4.8 | mean_score (max) | 31.25 |
| 8thof 72 | GPT-5.4 | mean_score (xhigh) | 27.1 |
| 9thof 72 | Claude Opus 4.7 | mean_score (xhigh) | 22.917 |
| 10thof 72 | Claude Opus 4.6 | mean_score (max) | 22.9 |
| 13thof 72 | GPT-5.2 | mean_score (xhigh) | 18.8 |
| 15thof 72 | gemini-3-pro-preview | mean_score | 18.75 |
| 16thof 72 | Gemini 3.1 Pro Preview | mean_score | 16.7 |
| 18thof 72 | Muse Spark 1.2 | mean_score | 14.6 |
| 19thof 72 | GPT-5 Pro | mean_score (high) | 14.6 |
| 20thof 72 | Gemini 3.5 Flash | mean_score (high) | 14.583 |
| 22ndof 72 | Kimi K2.6 | mean_score | 14.58 |
| 23rdof 72 | GLM 5.1 | mean_score | 12.5 |
| 24thof 72 | GPT-5.1 | mean_score (high) | 12.5 |
| 25thof 72 | GPT-5 | mean_score (high) | 12.5 |
| 27thof 72 | Qwen3.6 Plus | mean_score | 8.333 |
| 28thof 72 | Claude Sonnet 4.6 | mean_score | 8.3 |
| 29thof 72 | GPT-5.4 nano | mean_score (high) | 6.25 |
| 31stof 72 | GPT-5 mini | mean_score (high) | 6.25 |
| 33rdof 72 | o4-mini | mean_score (high) | 6.25 |
| 34thof 72 | Kimi K2.5 | mean_score | 4.2 |
| 35thof 72 | Qwen3.6 Max Preview | mean_score (max) | 4.167 |
| 36thof 72 | Gemini 2.5 Flash | mean_score | 4.167 |
| 37thof 72 | Gemini 3 Flash Preview | mean_score | 4.167 |
| 38thof 72 | Claude Opus 4.5 | mean_score | 4.167 |
| 41stof 72 | Claude Sonnet 4.5 | mean_score | 4.167 |
| 43rdof 72 | Claude Opus 4.1 | mean_score | 4.167 |
| 44thof 72 | Gemini 2.5 Pro | mean_score | 4.167 |
| 45thof 72 | o3-mini | mean_score (high) | 4.167 |
| 46thof 72 | Claude Opus 4 | mean_score | 4.167 |
| 47thof 72 | GLM-4.6 | mean_score | 2.128 |
| 48thof 72 | GLM-5 | mean_score | 2.1 |
| 49thof 72 | DeepSeek v3.2 | mean_score | 2.1 |
| 50thof 72 | Qwen 3.5 Plus | mean_score | 2.083 |
| 52ndof 72 | Claude Haiku 4.5 | mean_score | 2.083 |
| 56thof 72 | GPT-5 nano | mean_score (medium) | 2.083 |
| 57thof 72 | Gemini 2.5 Pro Preview 06-05 | mean_score | 2.083 |
| 58thof 72 | o3 | mean_score (high) | 2.083 |
| 59thof 72 | GPT-5.4 mini | mean_score (high) | 2.08 |
| 61stof 72 | Qwen3.6-Flash | mean_score | 0 |
| 62ndof 72 | Qwen3.5-Flash | mean_score | 0 |
| 63rdof 72 | GLM-4.7 | mean_score | 0 |
| 64thof 72 | Qwen3 235B A22B Thinking 2507 | mean_score | 0 |
| 65thof 72 | Kimi K2 Thinking | mean_score | 0 |
| 67thof 72 | GPT-4.1 fine-tuned | mean_score | 0 |
| 69thof 72 | Claude 3.5 Sonnet | mean_score | 0 |
| 71stof 72 | Grok 3 | mean_score | 0 |
| 72ndof 72 | Claude Sonnet 4 | mean_score | 0 |
Read from the board on 2026-09-16