GPQA Diamond
GPQA Diamond is run by Epoch AI (AI Benchmarking Hub). It has ranked 313 models, of which the catalogue holds 172, scoring from 9.28 to 95.77 on Accuracy.
measured by Epoch AI (AI Benchmarking Hub) · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 313 | GPT 6 Astra | mean_score (max) | 95.77 |
| 2ndof 313 | Gemini 3.8 Flash | mean_score (high) | 95.391 |
| 3rdof 313 | Gemini 3.7 Flash | mean_score (high) | 94.823 |
| 4thof 313 | GPT-5.4 Pro | mean_score (xhigh) | 94.6 |
| 5thof 313 | Gemini 3.1 Pro Preview | mean_score (high) | 94.444 |
| 6thof 313 | Gemini 3.6 Flash | mean_score (high) | 94.129 |
| 8thof 313 | Grok 4.6 | mean_score (high) | 94.003 |
| 9thof 313 | GPT-5.5 | mean_score (xhigh) | 94.003 |
| 10thof 313 | GPT-5.5 Pro | mean_score (xhigh) | 93.921 |
| 11thof 313 | Claude Opus 5 | mean_score (max) | 93.876 |
| 12thof 313 | GPT-5.6 Sol | mean_score (max) | 93.497 |
| 13thof 313 | Grok 4.5 | mean_score (high) | 93.434 |
| 14thof 313 | GPT-5.6 Terra | mean_score (max) | 93.308 |
| 15thof 313 | GPT-5.4 | mean_score (xhigh) | 93.3 |
| 17thof 313 | Kimi K3 | mean_score (max) | 93.119 |
| 19thof 313 | Gemini 3.5 Flash | mean_score (high) | 92.803 |
| 20thof 313 | Qwen 3.8 Max | mean_score (xhigh) | 92.677 |
| 21stof 313 | gemini-3-pro-preview | mean_score | 92.614 |
| 24thof 313 | GLM 5.2 | mean_score (max) | 91.856 |
| 25thof 313 | DeepSeek V4 Pro 0813 | mean_score (max) | 91.667 |
| 26thof 313 | GPT-5.6 Luna | mean_score (max) | 91.604 |
| 27thof 313 | GPT-5.2 | mean_score (xhigh) | 91.4 |
| 28thof 313 | DeepSeek V4 Flash (0731) | mean_score (max) | 91.035 |
| 29thof 313 | Claude Opus 4.8 | mean_score (max) | 91.035 |
| 30thof 313 | GLM-5.3 | mean_score (max) | 90.909 |
| 31stof 313 | MiniMax M3 | mean_score | 90.909 |
| 32ndof 313 | Qwen3.7 Max | mean_score (max) | 90.909 |
| 33rdof 313 | DeepSeek V4 Pro | mean_score (high) | 90.909 |
| 34thof 313 | Kimi K2.6 | mean_score | 90.783 |
| 36thof 313 | Claude Sonnet 5 | mean_score (xhigh) | 90.53 |
| 37thof 313 | Claude Opus 4.6 | mean_score | 90.53 |
| 38thof 313 | GLM 5.3 Flash | mean_score (max) | 90.152 |
| 39thof 313 | Claude Opus 4.7 | mean_score (xhigh) | 90.152 |
| 40thof 313 | GLM 5.1 | mean_score | 89.899 |
| 43rdof 313 | Muse Spark 1.2 | mean_score | 89.8 |
| 45thof 313 | Gemini 3 Flash | mean_score (high) | 89.394 |
| 46thof 313 | Grok 4.20 | mean_score | 89.331 |
| 49thof 313 | Grok 4.3 | mean_score (high) | 88.826 |
| 51stof 313 | Inkling Small | mean_score (xhigh) | 88.51 |
| 52ndof 313 | Qwen3.6 Plus | mean_score | 88.384 |
| 55thof 313 | Inkling | mean_score (xhigh) | 88.258 |
| 58thof 313 | Kimi K2.7 Code | mean_score | 87.879 |
| 59thof 313 | Qwen 3.7 Plus | mean_score | 87.879 |
| 62ndof 313 | GLM-5 | mean_score | 87.817 |
| 63rdof 313 | GPT-5.1 | mean_score (high) | 87.626 |
| 64thof 313 | Kimi K2.5 | mean_score | 87.6 |
| 65thof 313 | Qwen3.6 Max Preview | mean_score (max) | 87.374 |
| 67thof 313 | Claude Sonnet 4.6 | mean_score | 87.374 |
| 69thof 313 | GPT-5.4 mini | mean_score (xhigh) | 86.869 |
| 70thof 313 | Qwen3.5 397B A17B | mean_score (none) | 86.364 |
| 74thof 313 | GPT-5 | mean_score (high) | 86.174 |
| 75thof 313 | Claude Opus 4.5 | mean_score | 86.048 |
| 76thof 313 | Qwen3.6 27B | mean_score | 85.859 |
| 79thof 313 | Claude Fable 5 | mean_score (max) | 85.859 |
| 81stof 313 | Nemotron 3 Ultra | mean_score | 85.354 |
| 84thof 313 | Gemini 2.5 Pro | mean_score | 85.29 |
| 88thof 313 | Qwen 3.5 Plus | mean_score | 84.848 |
| 89thof 313 | Qwen3.6 35B A3B | mean_score (none) | 84.848 |
| 91stof 313 | Gemini 2.5 Pro Preview 06-05 | mean_score | 84.848 |
| 96thof 313 | Qwen3.5-35B-A3B | mean_score | 83.46 |
| 97thof 313 | deepseek-reasoner | mean_score | 83.424 |
| 98thof 313 | Qwen3.6-Flash | mean_score | 83.333 |
| 99thof 313 | Gemini 3.5 Flash-Lite | mean_score (high) | 83.333 |
| 103rdof 313 | GLM-4.7 | mean_score | 83.333 |
| 104thof 313 | Gemini 3 Flash Preview | mean_score | 83.207 |
| 107thof 313 | GPT-5.5 Instant | mean_score | 82.513 |
| 109thof 313 | Qwen3.7 Flash | mean_score | 82.323 |
| 110thof 313 | Qwen3.5-Flash | mean_score | 82.323 |
| 111thof 313 | Claude Sonnet 4.5 | mean_score | 82.323 |
| 113thof 313 | Gemini 3.1 Flash-Lite | mean_score (high) | 81.818 |
| 114thof 313 | o3 | mean_score (high) | 81.818 |
| 122ndof 313 | Qwen3 235B A22B Thinking 2507 | mean_score | 80.051 |
| 124thof 313 | o4-mini | mean_score (high) | 79.609 |
| 125thof 313 | Qwen3.5 9B | mean_score | 78.977 |
| 130thof 313 | Claude 3.7 Sonnet | mean_score | 78.504 |
| 131stof 313 | GPT-5.4 nano | mean_score (high) | 78.47 |
| 132ndof 313 | Claude Sonnet 4 | mean_score | 78.283 |
| 137thof 313 | Claude Opus 4.1 | mean_score | 77.273 |
| 138thof 313 | o3-mini | mean_score (high) | 77.02 |
| 142ndof 313 | o1 | mean_score (high) | 76.768 |
| 143rdof 313 | DeepSeek-R1 (0528) | mean_score | 76.326 |
| 144thof 313 | Claude Opus 4 | mean_score | 76.263 |
| 145thof 313 | Grok 3 Mini | mean_score (low) | 76.263 |
| 146thof 313 | Gemma 4 31B | mean_score (minimal) | 75.758 |
| 148thof 313 | GPT OSS 120B | mean_score (high) | 75.758 |
| 152ndof 313 | GPT-5 mini | mean_score (high) | 75 |
| 170thof 313 | Seed-OSS-36B-Instruct | mean_score | 71.528 |
| 172ndof 313 | deepseek-chat | mean_score | 71.212 |
| 173rdof 313 | Claude Haiku 4.5 | mean_score | 71.212 |
| 174thof 313 | Qwen3 235B A22B | mean_score | 70.707 |
| 175thof 313 | Qwen3 30B A3B Thinking 2507 | mean_score | 70.076 |
| 176thof 313 | GPT-5 nano | mean_score (high) | 69.444 |
| 177thof 313 | DeepSeek-R1 | mean_score | 69.223 |
| 181stof 313 | DeepSeek-V3 0324 | mean_score | 67.614 |
| 182ndof 313 | Grok 3 | mean_score | 67.582 |
| 184thof 313 | Llama 4 Maverick 17B | mean_score | 66.982 |
| 185thof 313 | GPT-4.1 fine-tuned | mean_score | 66.919 |
| 187thof 313 | Gemini 2.5 Pro Preview 05-06 | mean_score | 66.667 |
| 190thof 313 | GPT-4.1 mini | mean_score | 65.846 |
| 191stof 313 | Qwen3 32B | mean_score | 65.72 |
| 193rdof 313 | QwQ-Plus | mean_score | 65.404 |
| 194thof 313 | QwQ-32B | mean_score | 65.341 |
| 195thof 313 | DeepSeek R1 Distill QWEN 32B | mean_score | 64.141 |
| 197thof 313 | gemini-2.0-flash-001 | mean_score | 64.141 |
| 198thof 313 | Qwen3 14B | mean_score | 63.763 |
| 200thof 313 | o1-mini | mean_score (high) | 62.374 |
| 201stof 313 | Qwen3 30B A3B | mean_score | 61.679 |
| 202ndof 313 | GPT OSS 20B | mean_score (medium) | 60.795 |
| 203rdof 313 | GLM-4.7 Flash | mean_score | 60.543 |
| 205thof 313 | Mistral Medium 2505 | mean_score | 59.533 |
| 210thof 313 | Qwen3 8B | mean_score | 56.755 |
| 211thof 313 | DeepSeek-V3 | mean_score | 56.534 |
| 213thof 313 | Magistral Small 2506 | mean_score | 56.061 |
| 214thof 313 | Phi-4 | mean_score | 56.061 |
| 215thof 313 | R1 Distill Llama 70B | mean_score | 55.745 |
| 216thof 313 | Qwen3 30B A3B Instruct 2507 | mean_score | 55.619 |
| 218thof 313 | Claude 3.5 Sonnet | mean_score | 55.303 |
| 224thof 313 | Qwen3 4B | mean_score | 52.336 |
| 227thof 313 | Meta-Llama-3.1-405B-Instruct | mean_score | 50.915 |
| 230thof 313 | GPT-4o (2024-08-06) | mean_score | 49.211 |
| 231stof 313 | Qwen2.5 72B Instruct | mean_score | 49.148 |
| 233rdof 313 | Mistral Large 2407 | mean_score | 49.021 |
| 234thof 313 | GPT-4.1 nano | mean_score | 48.927 |
| 235thof 313 | GPT-4o (2024-05-13) | mean_score | 48.895 |
| 237thof 313 | Qwen-Plus | mean_score | 48.106 |
| 238thof 313 | GPT-4o (2024-11-20) | mean_score | 47.885 |
| 239thof 313 | Gemma 3 27B | mean_score | 47.727 |
| 242ndof 313 | mistral-small-2503 | mean_score | 47.475 |
| 243rdof 313 | Llama 3.3 70B Instruct | mean_score | 47.443 |
| 246thof 313 | Claude 3 Opus | mean_score | 47.159 |
| 247thof 313 | GPT-4 Turbo | mean_score | 46.591 |
| 249thof 313 | AgentRL | mean_score | 46.086 |
| 252ndof 313 | Qwen3-4B-Instruct-2507 | mean_score | 45.833 |
| 255thof 313 | DeepSeek R1 Distill QWEN 14B | mean_score | 44.697 |
| 256thof 313 | Meta-Llama-3.1-70B-Instruct | mean_score | 44.192 |
| 258thof 313 | WizardLM-2 8x22B | mean_score | 43.434 |
| 261stof 313 | Mistral Small 3.1 (25.03) | mean_score | 41.919 |
| 262ndof 313 | qwen-turbo | mean_score | 41.793 |
| 263rdof 313 | Llama-3.2-90B-Vision-Instruct | mean_score | 41.035 |
| 264thof 313 | Qwen2-72B-Instruct | mean_score | 40.783 |
| 265thof 313 | Claude 3 Sonnet | mean_score | 40.593 |
| 266thof 313 | Meta-Llama-3-70B-Instruct | mean_score | 40.562 |
| 268thof 313 | Gemma 3 12B | mean_score | 39.457 |
| 269thof 313 | Mistral Large (24.02) | mean_score | 38.763 |
| 270thof 313 | Claude Haiku 3.5 | mean_score | 38.131 |
| 271stof 313 | Qwen3-1.7B | mean_score | 38.005 |
| 272ndof 313 | GPT-4o-mini (2024-07-18) | mean_score | 37.721 |
| 274thof 313 | Gemma 2 27B | mean_score | 36.49 |
| 275thof 313 | Claude 3 Haiku | mean_score | 36.301 |
| 277thof 313 | Qwen2.5 7B Instruct | mean_score | 35.48 |
| 278thof 313 | Claude 2.0 | mean_score | 34.659 |
| 282ndof 313 | DeepSeek-R1-Distill-Qwen-1.5B | mean_score | 33.586 |
| 284thof 313 | Claude 2.1 | mean_score | 32.955 |
| 286thof 313 | Yi-1.5-34B-Chat | mean_score | 31.976 |
| 288thof 313 | GPT-4 (0613) | mean_score | 30.65 |
| 290thof 313 | Mixtral-8x7B-Instruct-v0.1 | mean_score | 30.587 |
| 293rdof 313 | Qwen1.5-72B-Chat | mean_score | 28.819 |
| 294thof 313 | Granite 4.0 Micro | mean_score | 28.283 |
| 295thof 313 | GPT-3.5 Turbo (1106) | mean_score | 28.03 |
| 296thof 313 | Phi-3-medium-128k-instruct | mean_score | 27.588 |
| 297thof 313 | gemma-2-9b-it | mean_score | 27.462 |
| 298thof 313 | GPT-3.5 Turbo | mean_score | 27.178 |
| 300thof 313 | Llama 3.1 8B Instruct | mean_score | 26.957 |
| 301stof 313 | Llama-2-70b-chat-hf | mean_score | 26.326 |
| 302ndof 313 | Llama 3 8B Instruct | mean_score | 26.073 |
| 304thof 313 | deepseek-llm-67b-chat | mean_score | 24.621 |
| 306thof 313 | Llama 3.2 1B Instruct | mean_score | 23.927 |
| 307thof 313 | Gemma 3 4B | mean_score | 23.232 |
| 309thof 313 | Mistral-7B-Instruct-v0.3 | mean_score | 15.183 |
| 310thof 313 | Yi-34B-Chat | mean_score | 14.741 |
| 311thof 313 | open-mistral-7b | mean_score | 13.226 |
| 313thof 313 | DeepSeek R1 0528 Qwen3 8B | mean_score | 9.28 |
Read from the board on 2026-09-16