Pass IndexThe State of AISign in

GPQA Diamond

GPQA Diamond is run by Epoch AI (AI Benchmarking Hub). It has ranked 313 models, of which the catalogue holds 172, scoring from 9.28 to 95.77 on Accuracy.

measured by Epoch AI (AI Benchmarking Hub) · the board itself

PlaceModelMetricScore
1stof 313GPT 6 Astramean_score (max)95.77
2ndof 313Gemini 3.8 Flashmean_score (high)95.391
3rdof 313Gemini 3.7 Flashmean_score (high)94.823
4thof 313GPT-5.4 Promean_score (xhigh)94.6
5thof 313Gemini 3.1 Pro Previewmean_score (high)94.444
6thof 313Gemini 3.6 Flashmean_score (high)94.129
8thof 313Grok 4.6mean_score (high)94.003
9thof 313GPT-5.5mean_score (xhigh)94.003
10thof 313GPT-5.5 Promean_score (xhigh)93.921
11thof 313Claude Opus 5mean_score (max)93.876
12thof 313GPT-5.6 Solmean_score (max)93.497
13thof 313Grok 4.5mean_score (high)93.434
14thof 313GPT-5.6 Terramean_score (max)93.308
15thof 313GPT-5.4mean_score (xhigh)93.3
17thof 313Kimi K3mean_score (max)93.119
19thof 313Gemini 3.5 Flashmean_score (high)92.803
20thof 313Qwen 3.8 Maxmean_score (xhigh)92.677
21stof 313gemini-3-pro-previewmean_score92.614
24thof 313GLM 5.2mean_score (max)91.856
25thof 313DeepSeek V4 Pro 0813mean_score (max)91.667
26thof 313GPT-5.6 Lunamean_score (max)91.604
27thof 313GPT-5.2mean_score (xhigh)91.4
28thof 313DeepSeek V4 Flash (0731)mean_score (max)91.035
29thof 313Claude Opus 4.8mean_score (max)91.035
30thof 313GLM-5.3mean_score (max)90.909
31stof 313MiniMax M3mean_score90.909
32ndof 313Qwen3.7 Maxmean_score (max)90.909
33rdof 313DeepSeek V4 Promean_score (high)90.909
34thof 313Kimi K2.6mean_score90.783
36thof 313Claude Sonnet 5mean_score (xhigh)90.53
37thof 313Claude Opus 4.6mean_score90.53
38thof 313GLM 5.3 Flashmean_score (max)90.152
39thof 313Claude Opus 4.7mean_score (xhigh)90.152
40thof 313GLM 5.1mean_score89.899
43rdof 313Muse Spark 1.2mean_score89.8
45thof 313Gemini 3 Flashmean_score (high)89.394
46thof 313Grok 4.20mean_score89.331
49thof 313Grok 4.3mean_score (high)88.826
51stof 313Inkling Smallmean_score (xhigh)88.51
52ndof 313Qwen3.6 Plusmean_score88.384
55thof 313Inklingmean_score (xhigh)88.258
58thof 313Kimi K2.7 Codemean_score87.879
59thof 313Qwen 3.7 Plusmean_score87.879
62ndof 313GLM-5mean_score87.817
63rdof 313GPT-5.1mean_score (high)87.626
64thof 313Kimi K2.5mean_score87.6
65thof 313Qwen3.6 Max Previewmean_score (max)87.374
67thof 313Claude Sonnet 4.6mean_score87.374
69thof 313GPT-5.4 minimean_score (xhigh)86.869
70thof 313Qwen3.5 397B A17Bmean_score (none)86.364
74thof 313GPT-5mean_score (high)86.174
75thof 313Claude Opus 4.5mean_score86.048
76thof 313Qwen3.6 27Bmean_score85.859
79thof 313Claude Fable 5mean_score (max)85.859
81stof 313Nemotron 3 Ultramean_score85.354
84thof 313Gemini 2.5 Promean_score85.29
88thof 313Qwen 3.5 Plusmean_score84.848
89thof 313Qwen3.6 35B A3Bmean_score (none)84.848
91stof 313Gemini 2.5 Pro Preview 06-05mean_score84.848
96thof 313Qwen3.5-35B-A3Bmean_score83.46
97thof 313deepseek-reasonermean_score83.424
98thof 313Qwen3.6-Flashmean_score83.333
99thof 313Gemini 3.5 Flash-Litemean_score (high)83.333
103rdof 313GLM-4.7mean_score83.333
104thof 313Gemini 3 Flash Previewmean_score83.207
107thof 313GPT-5.5 Instantmean_score82.513
109thof 313Qwen3.7 Flashmean_score82.323
110thof 313Qwen3.5-Flashmean_score82.323
111thof 313Claude Sonnet 4.5mean_score82.323
113thof 313Gemini 3.1 Flash-Litemean_score (high)81.818
114thof 313o3mean_score (high)81.818
122ndof 313Qwen3 235B A22B Thinking 2507mean_score80.051
124thof 313o4-minimean_score (high)79.609
125thof 313Qwen3.5 9Bmean_score78.977
130thof 313Claude 3.7 Sonnetmean_score78.504
131stof 313GPT-5.4 nanomean_score (high)78.47
132ndof 313Claude Sonnet 4mean_score78.283
137thof 313Claude Opus 4.1mean_score77.273
138thof 313o3-minimean_score (high)77.02
142ndof 313o1mean_score (high)76.768
143rdof 313DeepSeek-R1 (0528)mean_score76.326
144thof 313Claude Opus 4mean_score76.263
145thof 313Grok 3 Minimean_score (low)76.263
146thof 313Gemma 4 31Bmean_score (minimal)75.758
148thof 313GPT OSS 120Bmean_score (high)75.758
152ndof 313GPT-5 minimean_score (high)75
170thof 313Seed-OSS-36B-Instructmean_score71.528
172ndof 313deepseek-chatmean_score71.212
173rdof 313Claude Haiku 4.5mean_score71.212
174thof 313Qwen3 235B A22Bmean_score70.707
175thof 313Qwen3 30B A3B Thinking 2507mean_score70.076
176thof 313GPT-5 nanomean_score (high)69.444
177thof 313DeepSeek-R1mean_score69.223
181stof 313DeepSeek-V3 0324mean_score67.614
182ndof 313Grok 3mean_score67.582
184thof 313Llama 4 Maverick 17Bmean_score66.982
185thof 313GPT-4.1 fine-tunedmean_score66.919
187thof 313Gemini 2.5 Pro Preview 05-06mean_score66.667
190thof 313GPT-4.1 minimean_score65.846
191stof 313Qwen3 32Bmean_score65.72
193rdof 313QwQ-Plusmean_score65.404
194thof 313QwQ-32Bmean_score65.341
195thof 313DeepSeek R1 Distill QWEN 32Bmean_score64.141
197thof 313gemini-2.0-flash-001mean_score64.141
198thof 313Qwen3 14Bmean_score63.763
200thof 313o1-minimean_score (high)62.374
201stof 313Qwen3 30B A3Bmean_score61.679
202ndof 313GPT OSS 20Bmean_score (medium)60.795
203rdof 313GLM-4.7 Flashmean_score60.543
205thof 313Mistral Medium 2505mean_score59.533
210thof 313Qwen3 8Bmean_score56.755
211thof 313DeepSeek-V3mean_score56.534
213thof 313Magistral Small 2506mean_score56.061
214thof 313Phi-4mean_score56.061
215thof 313R1 Distill Llama 70Bmean_score55.745
216thof 313Qwen3 30B A3B Instruct 2507mean_score55.619
218thof 313Claude 3.5 Sonnetmean_score55.303
224thof 313Qwen3 4Bmean_score52.336
227thof 313Meta-Llama-3.1-405B-Instructmean_score50.915
230thof 313GPT-4o (2024-08-06)mean_score49.211
231stof 313Qwen2.5 72B Instructmean_score49.148
233rdof 313Mistral Large 2407mean_score49.021
234thof 313GPT-4.1 nanomean_score48.927
235thof 313GPT-4o (2024-05-13)mean_score48.895
237thof 313Qwen-Plusmean_score48.106
238thof 313GPT-4o (2024-11-20)mean_score47.885
239thof 313Gemma 3 27Bmean_score47.727
242ndof 313mistral-small-2503mean_score47.475
243rdof 313Llama 3.3 70B Instructmean_score47.443
246thof 313Claude 3 Opusmean_score47.159
247thof 313GPT-4 Turbomean_score46.591
249thof 313AgentRLmean_score46.086
252ndof 313Qwen3-4B-Instruct-2507mean_score45.833
255thof 313DeepSeek R1 Distill QWEN 14Bmean_score44.697
256thof 313Meta-Llama-3.1-70B-Instructmean_score44.192
258thof 313WizardLM-2 8x22Bmean_score43.434
261stof 313Mistral Small 3.1 (25.03)mean_score41.919
262ndof 313qwen-turbomean_score41.793
263rdof 313Llama-3.2-90B-Vision-Instructmean_score41.035
264thof 313Qwen2-72B-Instructmean_score40.783
265thof 313Claude 3 Sonnetmean_score40.593
266thof 313Meta-Llama-3-70B-Instructmean_score40.562
268thof 313Gemma 3 12Bmean_score39.457
269thof 313Mistral Large (24.02)mean_score38.763
270thof 313Claude Haiku 3.5mean_score38.131
271stof 313Qwen3-1.7Bmean_score38.005
272ndof 313GPT-4o-mini (2024-07-18)mean_score37.721
274thof 313Gemma 2 27Bmean_score36.49
275thof 313Claude 3 Haikumean_score36.301
277thof 313Qwen2.5 7B Instructmean_score35.48
278thof 313Claude 2.0mean_score34.659
282ndof 313DeepSeek-R1-Distill-Qwen-1.5Bmean_score33.586
284thof 313Claude 2.1mean_score32.955
286thof 313Yi-1.5-34B-Chatmean_score31.976
288thof 313GPT-4 (0613)mean_score30.65
290thof 313Mixtral-8x7B-Instruct-v0.1mean_score30.587
293rdof 313Qwen1.5-72B-Chatmean_score28.819
294thof 313Granite 4.0 Micromean_score28.283
295thof 313GPT-3.5 Turbo (1106)mean_score28.03
296thof 313Phi-3-medium-128k-instructmean_score27.588
297thof 313gemma-2-9b-itmean_score27.462
298thof 313GPT-3.5 Turbomean_score27.178
300thof 313Llama 3.1 8B Instructmean_score26.957
301stof 313Llama-2-70b-chat-hfmean_score26.326
302ndof 313Llama 3 8B Instructmean_score26.073
304thof 313deepseek-llm-67b-chatmean_score24.621
306thof 313Llama 3.2 1B Instructmean_score23.927
307thof 313Gemma 3 4Bmean_score23.232
309thof 313Mistral-7B-Instruct-v0.3mean_score15.183
310thof 313Yi-34B-Chatmean_score14.741
311thof 313open-mistral-7bmean_score13.226
313thof 313DeepSeek R1 0528 Qwen3 8Bmean_score9.28

Read from the board on 2026-09-16