OTIS Mock AIME 2024-2025
OTIS Mock AIME 2024-2025 is run by Epoch AI (AI Benchmarking Hub). It has ranked 290 models, of which the catalogue holds 160, scoring from 0 to 100 on Accuracy.
measured by Epoch AI (AI Benchmarking Hub) · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 2ndof 290 | Claude Fable 5.1 | mean_score (max) | 100 |
| 3rdof 290 | GPT 6 Astra | mean_score (max) | 100 |
| 4thof 290 | Claude Fable 5 | mean_score (high) | 100 |
| 5thof 290 | GPT-5.6 Sol | mean_score (max) | 100 |
| 6thof 290 | GPT-5.5 Pro | mean_score (xhigh) | 100 |
| 7thof 290 | GPT-5.5 | mean_score (xhigh) | 100 |
| 8thof 290 | GPT-5.6 Terra | mean_score (max) | 99.722 |
| 10thof 290 | Qwen 3.8 Max | mean_score (xhigh) | 99.444 |
| 11thof 290 | Grok 4.6 | mean_score (xhigh) | 99.167 |
| 12thof 290 | Gemini 3.8 Flash | mean_score (high) | 98.889 |
| 13thof 290 | Claude Opus 5 | mean_score (max) | 98.889 |
| 14thof 290 | DeepSeek V4 Pro 0813 | mean_score (max) | 98.611 |
| 15thof 290 | GPT-5.6 Luna | mean_score (max) | 98.333 |
| 16thof 290 | Claude Opus 4.8 | mean_score (max) | 98.333 |
| 17thof 290 | Claude Opus 4.7 | mean_score (xhigh) | 97.8 |
| 22ndof 290 | GPT-5.4 | mean_score (high) | 97.778 |
| 23rdof 290 | Grok 4.5 | mean_score (high) | 97.778 |
| 24thof 290 | Gemini 3.7 Flash | mean_score (high) | 97.222 |
| 25thof 290 | Kimi K3 | mean_score (max) | 97.222 |
| 26thof 290 | DeepSeek V4 Pro | mean_score (max) | 96.667 |
| 27thof 290 | Kimi K2.6 | mean_score | 96.111 |
| 28thof 290 | GPT-5.2 | mean_score (high) | 96.111 |
| 30thof 290 | Gemini 3.1 Pro Preview | mean_score | 95.6 |
| 32ndof 290 | Qwen3.7 Max | mean_score (max) | 95.556 |
| 33rdof 290 | Kimi K2.7 Code | mean_score | 95.556 |
| 36thof 290 | Gemini 3 Flash | mean_score (high) | 95.556 |
| 38thof 290 | Gemini 3.5 Flash | mean_score (high) | 95.556 |
| 40thof 290 | Claude Sonnet 5 | mean_score (xhigh) | 94.722 |
| 41stof 290 | DeepSeek V4 Flash (0731) | mean_score (max) | 94.444 |
| 42ndof 290 | Claude Opus 4.6 | mean_score | 94.444 |
| 43rdof 290 | Gemini 3.6 Flash | mean_score (high) | 94.167 |
| 44thof 290 | GLM 5.3 Flash | mean_score (max) | 93.889 |
| 46thof 290 | GLM 5.1 | mean_score | 93.333 |
| 48thof 290 | Qwen3.6 Plus | mean_score | 93.333 |
| 49thof 290 | Qwen 3.7 Plus | mean_score | 93.333 |
| 51stof 290 | Grok 4.3 | mean_score (high) | 93.333 |
| 53rdof 290 | Gemini 3 Flash Preview | mean_score | 92.778 |
| 54thof 290 | Grok 4.20 | mean_score | 92.222 |
| 55thof 290 | Kimi K2.5 | mean_score | 92.2 |
| 56thof 290 | gemini-3-pro-preview | mean_score | 91.389 |
| 57thof 290 | GPT-5 | mean_score (high) | 91.389 |
| 58thof 290 | GLM-5.3 | mean_score (max) | 91.111 |
| 59thof 290 | Qwen3.6 Max Preview | mean_score (max) | 91.111 |
| 60thof 290 | Qwen3.6 27B | mean_score | 91.111 |
| 62ndof 290 | Inkling Small | mean_score (xhigh) | 90 |
| 63rdof 290 | Muse Spark 1.2 | mean_score | 88.9 |
| 64thof 290 | GPT-5.4 mini | mean_score (xhigh) | 88.889 |
| 66thof 290 | Qwen3.5 397B A17B | mean_score | 88.889 |
| 68thof 290 | Inkling | mean_score (xhigh) | 88.889 |
| 69thof 290 | GPT OSS 120B | mean_score (high) | 88.889 |
| 70thof 290 | GPT-5.1 | mean_score (high) | 88.611 |
| 71stof 290 | deepseek-reasoner | mean_score | 87.817 |
| 72ndof 290 | GPT-5.4 nano | mean_score (high) | 87.78 |
| 75thof 290 | Nemotron 3 Ultra | mean_score | 86.667 |
| 76thof 290 | Qwen 3.5 Plus | mean_score | 86.667 |
| 77thof 290 | Qwen3.6 35B A3B | mean_score | 86.667 |
| 78thof 290 | Qwen3.7 Flash | mean_score | 86.667 |
| 80thof 290 | Qwen3 235B A22B Thinking 2507 | mean_score | 86.667 |
| 81stof 290 | GPT-5 mini | mean_score (high) | 86.667 |
| 82ndof 290 | GLM 5.2 | mean_score (max) | 86.389 |
| 83rdof 290 | Claude Opus 4.5 | mean_score | 86.111 |
| 84thof 290 | Claude Sonnet 4.6 | mean_score | 85.794 |
| 87thof 290 | o3 | mean_score (medium) | 84.444 |
| 88thof 290 | Qwen3.6-Flash | mean_score | 84.444 |
| 89thof 290 | Qwen3.5-Flash | mean_score | 84.444 |
| 92ndof 290 | Gemini 2.5 Pro | mean_score | 84.167 |
| 95thof 290 | GLM-4.7 | mean_score | 83.333 |
| 102ndof 290 | o4-mini | mean_score (high) | 81.667 |
| 103rdof 290 | GPT-5 nano | mean_score (high) | 81.111 |
| 106thof 290 | Gemini 3.1 Flash-Lite | mean_score (high) | 80 |
| 109thof 290 | GLM-5 | mean_score | 80 |
| 113thof 290 | Claude Sonnet 4.5 | mean_score | 77.778 |
| 115thof 290 | Grok 3 Mini | mean_score (high) | 77.778 |
| 116thof 290 | o3-mini | mean_score (high) | 76.944 |
| 121stof 290 | Gemma 4 31B | mean_score (minimal) | 73.333 |
| 122ndof 290 | o1 | mean_score (high) | 73.333 |
| 125thof 290 | Gemini 2.5 Flash | mean_score | 73.056 |
| 126thof 290 | MiniMax M3 | mean_score | 71.111 |
| 127thof 290 | Gemini 3.5 Flash-Lite | mean_score (high) | 71.111 |
| 130thof 290 | Claude Sonnet 4 | mean_score | 71.111 |
| 132ndof 290 | Qwen3 30B A3B Thinking 2507 | mean_score | 70.278 |
| 133rdof 290 | Qwen3.5-35B-A3B | mean_score (none) | 70 |
| 138thof 290 | Claude Opus 4.1 | mean_score | 68.889 |
| 140thof 290 | GPT-5.5 Instant | mean_score | 68.056 |
| 141stof 290 | Seed-OSS-36B-Instruct | mean_score | 67.5 |
| 142ndof 290 | Qwen3 32B | mean_score | 66.944 |
| 145thof 290 | Claude Haiku 4.5 | mean_score | 66.667 |
| 146thof 290 | Qwen3 14B | mean_score | 66.389 |
| 147thof 290 | DeepSeek-R1 (0528) | mean_score | 66.389 |
| 148thof 290 | GPT OSS 20B | mean_score (medium) | 65.278 |
| 150thof 290 | Claude Opus 4 | mean_score | 64.444 |
| 153rdof 290 | Qwen3 30B A3B | mean_score | 62.778 |
| 154thof 290 | Qwen3 30B A3B Instruct 2507 | mean_score | 62.222 |
| 157thof 290 | Qwen3.5 9B | mean_score (none) | 61.667 |
| 161stof 290 | QwQ-32B | mean_score | 59.167 |
| 162ndof 290 | GLM-4.7 Flash | mean_score | 58.333 |
| 166thof 290 | Claude 3.7 Sonnet | mean_score | 57.778 |
| 168thof 290 | Qwen3 8B | mean_score | 56.111 |
| 169thof 290 | Qwen3.5-4B | mean_score (none) | 55.833 |
| 170thof 290 | DeepSeek R1 Distill QWEN 32B | mean_score | 55.556 |
| 172ndof 290 | Grok 3 | mean_score | 55.556 |
| 178thof 290 | DeepSeek-R1 | mean_score | 53.333 |
| 179thof 290 | Qwen3-4B-Instruct-2507 | mean_score | 52.222 |
| 180thof 290 | R1 Distill Llama 70B | mean_score | 51.389 |
| 183rdof 290 | DeepSeek R1 Distill QWEN 14B | mean_score | 50.556 |
| 184thof 290 | deepseek-chat | mean_score | 48.889 |
| 186thof 290 | o1-mini | mean_score (high) | 46.944 |
| 192ndof 290 | GPT-4.1 mini | mean_score | 44.722 |
| 196thof 290 | DeepSeek R1 0528 Qwen3 8B | mean_score | 43.889 |
| 202ndof 290 | GPT-4.1 fine-tuned | mean_score | 38.333 |
| 205thof 290 | DeepSeek-V3 0324 | mean_score | 37.778 |
| 210thof 290 | Mistral Medium 2505 | mean_score | 32.222 |
| 212thof 290 | gemini-2.0-flash-001 | mean_score | 31.111 |
| 214thof 290 | Magistral Small 2506 | mean_score | 30 |
| 217thof 290 | GPT-4.1 nano | mean_score | 28.889 |
| 226thof 290 | Gemma 3 27B | mean_score | 22.5 |
| 229thof 290 | DeepSeek-R1-Distill-Qwen-1.5B | mean_score | 21.389 |
| 230thof 290 | Llama 4 Maverick 17B | mean_score | 20.556 |
| 231stof 290 | Qwen-Plus | mean_score | 17.778 |
| 232ndof 290 | Gemma 3 12B | mean_score | 16.667 |
| 235thof 290 | DeepSeek-V3 | mean_score | 15.833 |
| 236thof 290 | Phi-4 | mean_score | 13.75 |
| 238thof 290 | Meta-Llama-3.1-405B-Instruct | mean_score | 9.722 |
| 239thof 290 | Claude 3.5 Sonnet | mean_score | 8.472 |
| 240thof 290 | Mistral Large 2407 | mean_score | 8.472 |
| 241stof 290 | Qwen3-1.7B | mean_score (none) | 8.056 |
| 242ndof 290 | Qwen2.5 72B Instruct | mean_score | 8.056 |
| 245thof 290 | Gemma 3 4B | mean_score | 7.5 |
| 246thof 290 | AgentRL | mean_score | 7.361 |
| 247thof 290 | GPT-4o-mini (2024-07-18) | mean_score | 6.944 |
| 250thof 290 | GPT-4 Turbo | mean_score | 6.667 |
| 252ndof 290 | GPT-4o (2024-08-06) | mean_score | 6.389 |
| 253rdof 290 | GPT-4o (2024-05-13) | mean_score | 6.25 |
| 254thof 290 | GPT-4o (2024-11-20) | mean_score | 6.25 |
| 255thof 290 | qwen-turbo | mean_score | 6.111 |
| 256thof 290 | mistral-small-2503 | mean_score | 5.833 |
| 258thof 290 | Llama 3.3 70B Instruct | mean_score | 5.139 |
| 259thof 290 | Claude 3 Opus | mean_score | 4.722 |
| 262ndof 290 | Meta-Llama-3-70B-Instruct | mean_score | 4.306 |
| 263rdof 290 | Claude Haiku 3.5 | mean_score | 4.306 |
| 264thof 290 | Mistral Small 3.1 (25.03) | mean_score | 3.889 |
| 266thof 290 | Meta-Llama-3.1-70B-Instruct | mean_score | 3.611 |
| 267thof 290 | Granite 4.0 Micro | mean_score | 2.778 |
| 268thof 290 | Llama-3.2-90B-Vision-Instruct | mean_score | 2.639 |
| 269thof 290 | Qwen2.5 7B Instruct | mean_score | 2.5 |
| 271stof 290 | Claude 2.0 | mean_score | 2.5 |
| 272ndof 290 | Claude 3 Sonnet | mean_score | 2.5 |
| 273rdof 290 | GPT-3.5 Turbo | mean_score | 2.222 |
| 274thof 290 | Llama 3 8B Instruct | mean_score | 1.944 |
| 275thof 290 | Claude 2.1 | mean_score | 1.944 |
| 276thof 290 | Mistral Large (24.02) | mean_score | 1.944 |
| 277thof 290 | Claude 3 Haiku | mean_score | 1.806 |
| 278thof 290 | Llama 3.1 8B Instruct | mean_score | 1.667 |
| 279thof 290 | Gemma 2 27B | mean_score | 1.389 |
| 281stof 290 | GPT-4 (0613) | mean_score | 1.111 |
| 284thof 290 | deepseek-llm-67b-chat | mean_score | 83.3 |
| 285thof 290 | Llama 3.2 1B Instruct | mean_score | 55.6 |
| 287thof 290 | gemma-2-9b-it | mean_score | 55.6 |
| 288thof 290 | Mistral-7B-Instruct-v0.3 | mean_score | 27.8 |
| 290thof 290 | Llama-2-70b-chat-hf | mean_score | 0 |
Read from the board on 2026-09-16