Simpleqa verified — Epoch AI
Simpleqa verified — Epoch AI is run by Epoch AI. It has ranked 80 models, of which the catalogue holds 69, scoring from 6 to 75.6.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 80 | GPT 6 Astra | mean_score (max) | 75.6 |
| 2ndof 80 | Gemini 3.1 Pro Preview | mean_score (high) | 73.5 |
| 3rdof 80 | Claude Fable 5.1 | mean_score (max) | 70.8 |
| 4thof 80 | Claude Fable 5 | mean_score (xhigh) | 70.7 |
| 5thof 80 | Gemini 3.8 Flash | mean_score (high) | 69.7 |
| 6thof 80 | GPT-5.6 Sol | mean_score (max) | 69.7 |
| 7thof 80 | Gemini 3.7 Flash | mean_score (high) | 69.2 |
| 8thof 80 | Gemini 3 Flash | mean_score (high) | 66.8 |
| 9thof 80 | Gemini 3.5 Flash | mean_score (high) | 66.2 |
| 10thof 80 | Gemini 3.6 Flash | mean_score (high) | 66.2 |
| 11thof 80 | GPT-5.5 | mean_score (xhigh) | 63 |
| 12thof 80 | Muse Spark 1.2 | mean_score (xhigh) | 60.302 |
| 13thof 80 | Claude Opus 5 | mean_score (max) | 59.9 |
| 14thof 80 | Muse Spark 1.1 | mean_score | 57.789 |
| 15thof 80 | Qwen3.7 Max | mean_score (max) | 55.767 |
| 16thof 80 | Claude Opus 4.8 | mean_score (max) | 53 |
| 17thof 80 | DeepSeek V4 Pro 0813 | mean_score (max) | 52.906 |
| 18thof 80 | Qwen3.6 Max Preview | mean_score (max) | 52.004 |
| 19thof 80 | Claude Opus 4.7 | mean_score (xhigh) | 51.7 |
| 20thof 80 | Kimi K3 | mean_score (max) | 50.6 |
| 21stof 80 | GPT-5 | mean_score (high) | 50.1 |
| 22ndof 80 | o3 | mean_score (high) | 49.4 |
| 23rdof 80 | Grok 4.6 | mean_score (high) | 49.3 |
| 26thof 80 | Grok 4.5 | mean_score (high) | 48.3 |
| 27thof 80 | GPT-5.1 | mean_score (high) | 48 |
| 29thof 80 | Claude Opus 4.6 | mean_score (max) | 47 |
| 30thof 80 | DeepSeek V4 Pro | mean_score (max) | 46.994 |
| 31stof 80 | GPT-5.4 Pro | mean_score (xhigh) | 46.3 |
| 32ndof 80 | Qwen 3.8 Max | mean_score (xhigh) | 45.8 |
| 33rdof 80 | Claude Opus 4.5 | mean_score | 45.7 |
| 34thof 80 | GPT-5.4 | mean_score (xhigh) | 45.1 |
| 35thof 80 | Qwen3.6 Plus | mean_score | 44.144 |
| 36thof 80 | GPT-5.6 Terra | mean_score (max) | 43.2 |
| 37thof 80 | o1 | mean_score (high) | 41.1 |
| 38thof 80 | GLM-5.3 | mean_score (max) | 41 |
| 39thof 80 | GPT-5.6 Luna | mean_score (max) | 41 |
| 40thof 80 | Qwen3 235B A22B Thinking 2507 | mean_score | 40.44 |
| 41stof 80 | Inkling | mean_score (xhigh) | 40.3 |
| 42ndof 80 | GPT-5.2 | mean_score (xhigh) | 37.1 |
| 43rdof 80 | Kimi K2.7 Code | mean_score | 36.5 |
| 44thof 80 | Claude Sonnet 4.6 | mean_score (high) | 35.5 |
| 45thof 80 | Kimi K2.6 | mean_score | 34.9 |
| 46thof 80 | Kimi K2.5 | mean_score | 34.3 |
| 48thof 80 | GLM 5.2 | mean_score (max) | 34.2 |
| 49thof 80 | GLM 5.1 | mean_score | 34 |
| 50thof 80 | Claude Sonnet 5 | mean_score (max) | 33.7 |
| 51stof 80 | DeepSeek V4 Flash (0731) | mean_score (max) | 33.634 |
| 52ndof 80 | Grok 4.3 | mean_score (high) | 33.2 |
| 57thof 80 | GLM-4.7 | mean_score | 32.2 |
| 58thof 80 | GPT-4.1 fine-tuned | mean_score | 31.1 |
| 59thof 80 | Claude Sonnet 4.5 | mean_score | 30.7 |
| 60thof 80 | Grok 4.20 | mean_score | 30.2 |
| 61stof 80 | GPT-5.4 mini | mean_score (high) | 29.4 |
| 62ndof 80 | GPT-4o (2024-08-06) | mean_score | 26 |
| 63rdof 80 | Qwen 3.5 Plus | mean_score | 25.351 |
| 65thof 80 | GPT-5 mini | mean_score (high) | 21.6 |
| 66thof 80 | Qwen3.5-Flash | mean_score | 20.32 |
| 67thof 80 | o4-mini | mean_score (low) | 19.6 |
| 68thof 80 | Inkling Small | mean_score (xhigh) | 19.1 |
| 70thof 80 | Qwen3.6-Flash | mean_score | 15.916 |
| 71stof 80 | o3-mini | mean_score (high) | 15.3 |
| 72ndof 80 | Claude Haiku 4.5 | mean_score | 13.2 |
| 73rdof 80 | GPT-4.1 mini | mean_score | 12.7 |
| 74thof 80 | Claude 3 Opus | mean_score | 12.6 |
| 76thof 80 | GPT-5.4 nano | mean_score (high) | 11.7 |
| 77thof 80 | GPT-5 nano | mean_score (high) | 11.7 |
| 78thof 80 | Gemma 4 31B | mean_score | 10.4 |
| 79thof 80 | GPT-4o-mini (2024-07-18) | mean_score | 8.3 |
| 80thof 80 | GPT-4.1 nano | mean_score | 6 |
Read from the board on 2026-09-16