Pass IndexThe State of AISign in

OTIS Mock AIME 2024-2025

OTIS Mock AIME 2024-2025 is run by Epoch AI (AI Benchmarking Hub). It has ranked 290 models, of which the catalogue holds 160, scoring from 0 to 100 on Accuracy.

measured by Epoch AI (AI Benchmarking Hub) · the board itself

PlaceModelMetricScore
2ndof 290Claude Fable 5.1mean_score (max)100
3rdof 290GPT 6 Astramean_score (max)100
4thof 290Claude Fable 5mean_score (high)100
5thof 290GPT-5.6 Solmean_score (max)100
6thof 290GPT-5.5 Promean_score (xhigh)100
7thof 290GPT-5.5mean_score (xhigh)100
8thof 290GPT-5.6 Terramean_score (max)99.722
10thof 290Qwen 3.8 Maxmean_score (xhigh)99.444
11thof 290Grok 4.6mean_score (xhigh)99.167
12thof 290Gemini 3.8 Flashmean_score (high)98.889
13thof 290Claude Opus 5mean_score (max)98.889
14thof 290DeepSeek V4 Pro 0813mean_score (max)98.611
15thof 290GPT-5.6 Lunamean_score (max)98.333
16thof 290Claude Opus 4.8mean_score (max)98.333
17thof 290Claude Opus 4.7mean_score (xhigh)97.8
22ndof 290GPT-5.4mean_score (high)97.778
23rdof 290Grok 4.5mean_score (high)97.778
24thof 290Gemini 3.7 Flashmean_score (high)97.222
25thof 290Kimi K3mean_score (max)97.222
26thof 290DeepSeek V4 Promean_score (max)96.667
27thof 290Kimi K2.6mean_score96.111
28thof 290GPT-5.2mean_score (high)96.111
30thof 290Gemini 3.1 Pro Previewmean_score95.6
32ndof 290Qwen3.7 Maxmean_score (max)95.556
33rdof 290Kimi K2.7 Codemean_score95.556
36thof 290Gemini 3 Flashmean_score (high)95.556
38thof 290Gemini 3.5 Flashmean_score (high)95.556
40thof 290Claude Sonnet 5mean_score (xhigh)94.722
41stof 290DeepSeek V4 Flash (0731)mean_score (max)94.444
42ndof 290Claude Opus 4.6mean_score94.444
43rdof 290Gemini 3.6 Flashmean_score (high)94.167
44thof 290GLM 5.3 Flashmean_score (max)93.889
46thof 290GLM 5.1mean_score93.333
48thof 290Qwen3.6 Plusmean_score93.333
49thof 290Qwen 3.7 Plusmean_score93.333
51stof 290Grok 4.3mean_score (high)93.333
53rdof 290Gemini 3 Flash Previewmean_score92.778
54thof 290Grok 4.20mean_score92.222
55thof 290Kimi K2.5mean_score92.2
56thof 290gemini-3-pro-previewmean_score91.389
57thof 290GPT-5mean_score (high)91.389
58thof 290GLM-5.3mean_score (max)91.111
59thof 290Qwen3.6 Max Previewmean_score (max)91.111
60thof 290Qwen3.6 27Bmean_score91.111
62ndof 290Inkling Smallmean_score (xhigh)90
63rdof 290Muse Spark 1.2mean_score88.9
64thof 290GPT-5.4 minimean_score (xhigh)88.889
66thof 290Qwen3.5 397B A17Bmean_score88.889
68thof 290Inklingmean_score (xhigh)88.889
69thof 290GPT OSS 120Bmean_score (high)88.889
70thof 290GPT-5.1mean_score (high)88.611
71stof 290deepseek-reasonermean_score87.817
72ndof 290GPT-5.4 nanomean_score (high)87.78
75thof 290Nemotron 3 Ultramean_score86.667
76thof 290Qwen 3.5 Plusmean_score86.667
77thof 290Qwen3.6 35B A3Bmean_score86.667
78thof 290Qwen3.7 Flashmean_score86.667
80thof 290Qwen3 235B A22B Thinking 2507mean_score86.667
81stof 290GPT-5 minimean_score (high)86.667
82ndof 290GLM 5.2mean_score (max)86.389
83rdof 290Claude Opus 4.5mean_score86.111
84thof 290Claude Sonnet 4.6mean_score85.794
87thof 290o3mean_score (medium)84.444
88thof 290Qwen3.6-Flashmean_score84.444
89thof 290Qwen3.5-Flashmean_score84.444
92ndof 290Gemini 2.5 Promean_score84.167
95thof 290GLM-4.7mean_score83.333
102ndof 290o4-minimean_score (high)81.667
103rdof 290GPT-5 nanomean_score (high)81.111
106thof 290Gemini 3.1 Flash-Litemean_score (high)80
109thof 290GLM-5mean_score80
113thof 290Claude Sonnet 4.5mean_score77.778
115thof 290Grok 3 Minimean_score (high)77.778
116thof 290o3-minimean_score (high)76.944
121stof 290Gemma 4 31Bmean_score (minimal)73.333
122ndof 290o1mean_score (high)73.333
125thof 290Gemini 2.5 Flashmean_score73.056
126thof 290MiniMax M3mean_score71.111
127thof 290Gemini 3.5 Flash-Litemean_score (high)71.111
130thof 290Claude Sonnet 4mean_score71.111
132ndof 290Qwen3 30B A3B Thinking 2507mean_score70.278
133rdof 290Qwen3.5-35B-A3Bmean_score (none)70
138thof 290Claude Opus 4.1mean_score68.889
140thof 290GPT-5.5 Instantmean_score68.056
141stof 290Seed-OSS-36B-Instructmean_score67.5
142ndof 290Qwen3 32Bmean_score66.944
145thof 290Claude Haiku 4.5mean_score66.667
146thof 290Qwen3 14Bmean_score66.389
147thof 290DeepSeek-R1 (0528)mean_score66.389
148thof 290GPT OSS 20Bmean_score (medium)65.278
150thof 290Claude Opus 4mean_score64.444
153rdof 290Qwen3 30B A3Bmean_score62.778
154thof 290Qwen3 30B A3B Instruct 2507mean_score62.222
157thof 290Qwen3.5 9Bmean_score (none)61.667
161stof 290QwQ-32Bmean_score59.167
162ndof 290GLM-4.7 Flashmean_score58.333
166thof 290Claude 3.7 Sonnetmean_score57.778
168thof 290Qwen3 8Bmean_score56.111
169thof 290Qwen3.5-4Bmean_score (none)55.833
170thof 290DeepSeek R1 Distill QWEN 32Bmean_score55.556
172ndof 290Grok 3mean_score55.556
178thof 290DeepSeek-R1mean_score53.333
179thof 290Qwen3-4B-Instruct-2507mean_score52.222
180thof 290R1 Distill Llama 70Bmean_score51.389
183rdof 290DeepSeek R1 Distill QWEN 14Bmean_score50.556
184thof 290deepseek-chatmean_score48.889
186thof 290o1-minimean_score (high)46.944
192ndof 290GPT-4.1 minimean_score44.722
196thof 290DeepSeek R1 0528 Qwen3 8Bmean_score43.889
202ndof 290GPT-4.1 fine-tunedmean_score38.333
205thof 290DeepSeek-V3 0324mean_score37.778
210thof 290Mistral Medium 2505mean_score32.222
212thof 290gemini-2.0-flash-001mean_score31.111
214thof 290Magistral Small 2506mean_score30
217thof 290GPT-4.1 nanomean_score28.889
226thof 290Gemma 3 27Bmean_score22.5
229thof 290DeepSeek-R1-Distill-Qwen-1.5Bmean_score21.389
230thof 290Llama 4 Maverick 17Bmean_score20.556
231stof 290Qwen-Plusmean_score17.778
232ndof 290Gemma 3 12Bmean_score16.667
235thof 290DeepSeek-V3mean_score15.833
236thof 290Phi-4mean_score13.75
238thof 290Meta-Llama-3.1-405B-Instructmean_score9.722
239thof 290Claude 3.5 Sonnetmean_score8.472
240thof 290Mistral Large 2407mean_score8.472
241stof 290Qwen3-1.7Bmean_score (none)8.056
242ndof 290Qwen2.5 72B Instructmean_score8.056
245thof 290Gemma 3 4Bmean_score7.5
246thof 290AgentRLmean_score7.361
247thof 290GPT-4o-mini (2024-07-18)mean_score6.944
250thof 290GPT-4 Turbomean_score6.667
252ndof 290GPT-4o (2024-08-06)mean_score6.389
253rdof 290GPT-4o (2024-05-13)mean_score6.25
254thof 290GPT-4o (2024-11-20)mean_score6.25
255thof 290qwen-turbomean_score6.111
256thof 290mistral-small-2503mean_score5.833
258thof 290Llama 3.3 70B Instructmean_score5.139
259thof 290Claude 3 Opusmean_score4.722
262ndof 290Meta-Llama-3-70B-Instructmean_score4.306
263rdof 290Claude Haiku 3.5mean_score4.306
264thof 290Mistral Small 3.1 (25.03)mean_score3.889
266thof 290Meta-Llama-3.1-70B-Instructmean_score3.611
267thof 290Granite 4.0 Micromean_score2.778
268thof 290Llama-3.2-90B-Vision-Instructmean_score2.639
269thof 290Qwen2.5 7B Instructmean_score2.5
271stof 290Claude 2.0mean_score2.5
272ndof 290Claude 3 Sonnetmean_score2.5
273rdof 290GPT-3.5 Turbomean_score2.222
274thof 290Llama 3 8B Instructmean_score1.944
275thof 290Claude 2.1mean_score1.944
276thof 290Mistral Large (24.02)mean_score1.944
277thof 290Claude 3 Haikumean_score1.806
278thof 290Llama 3.1 8B Instructmean_score1.667
279thof 290Gemma 2 27Bmean_score1.389
281stof 290GPT-4 (0613)mean_score1.111
284thof 290deepseek-llm-67b-chatmean_score83.3
285thof 290Llama 3.2 1B Instructmean_score55.6
287thof 290gemma-2-9b-itmean_score55.6
288thof 290Mistral-7B-Instruct-v0.3mean_score27.8
290thof 290Llama-2-70b-chat-hfmean_score0

Read from the board on 2026-09-16