Chess puzzles — Epoch AI
Chess puzzles — Epoch AI is run by Epoch AI. It has ranked 222 models, of which the catalogue holds 125, scoring from 0 to 100.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 222 | GPT 6 Astra | mean_score (max) | 72 |
| 3rdof 222 | GPT-5.5 Pro | mean_score (xhigh) | 64 |
| 4thof 222 | Gemini 3.8 Flash | mean_score (high) | 61 |
| 5thof 222 | GPT-5.4 Pro | mean_score (xhigh) | 58.6 |
| 6thof 222 | GPT-5.6 Sol | mean_score (max) | 55 |
| 7thof 222 | Gemini 3.1 Pro Preview | mean_score | 55 |
| 8thof 222 | GPT-5.6 Terra | mean_score (max) | 54 |
| 9thof 222 | GPT-5.5 | mean_score (xhigh) | 54 |
| 10thof 222 | Gemini 3.5 Flash | mean_score (high) | 50 |
| 12thof 222 | GPT-5.2 | mean_score (xhigh) | 49 |
| 13thof 222 | Claude Fable 5.1 | mean_score (max) | 47 |
| 14thof 222 | DeepSeek V4 Pro 0813 | mean_score (max) | 47 |
| 15thof 222 | Gemini 3.7 Flash | mean_score (high) | 47 |
| 17thof 222 | GPT-5.4 | mean_score (xhigh) | 44 |
| 18thof 222 | Gemini 3.6 Flash | mean_score (low) | 43 |
| 20thof 222 | Claude Opus 5 | mean_score (max) | 42 |
| 21stof 222 | Claude Fable 5 | mean_score (high) | 41 |
| 24thof 222 | Grok 4.6 | mean_score (high) | 40 |
| 25thof 222 | Gemini 3 Flash | mean_score (high) | 40 |
| 27thof 222 | GPT-5.6 Luna | mean_score (max) | 40 |
| 30thof 222 | Kimi K3 | mean_score (max) | 39 |
| 31stof 222 | o3 | mean_score (medium) | 38 |
| 34thof 222 | Gemini 3 Flash Preview | mean_score | 38 |
| 35thof 222 | GPT-5 | mean_score (high) | 37 |
| 36thof 222 | Grok 4.5 | mean_score (high) | 36 |
| 38thof 222 | Claude Sonnet 5 | mean_score (xhigh) | 35 |
| 40thof 222 | Claude Opus 4.8 | mean_score (max) | 34 |
| 42ndof 222 | DeepSeek V4 Flash (0731) | mean_score (max) | 33 |
| 43rdof 222 | GPT-5.1 | mean_score (high) | 32 |
| 45thof 222 | gemini-3-pro-preview | mean_score | 31 |
| 46thof 222 | GPT-5 mini | mean_score (high) | 30 |
| 47thof 222 | Claude Opus 4.7 | mean_score (xhigh) | 30 |
| 48thof 222 | GPT-5.4 nano | mean_score (high) | 30 |
| 50thof 222 | Qwen 3.8 Max | mean_score (xhigh) | 29 |
| 55thof 222 | GPT-5 nano | mean_score (high) | 27 |
| 58thof 222 | Qwen3.6 35B A3B | mean_score | 26 |
| 59thof 222 | Kimi K2.6 | mean_score | 26 |
| 60thof 222 | o4-mini | mean_score (high) | 26 |
| 62ndof 222 | Gemini 3.1 Flash-Lite | mean_score (low) | 25 |
| 63rdof 222 | Grok 4.3 | mean_score (high) | 25 |
| 64thof 222 | GPT-5.4 mini | mean_score (xhigh) | 24 |
| 65thof 222 | Qwen 3.7 Plus | mean_score | 24 |
| 68thof 222 | Grok 4.20 | mean_score | 24 |
| 69thof 222 | Qwen3.7 Flash | mean_score | 23 |
| 72ndof 222 | Qwen3.6 27B | mean_score | 22 |
| 73rdof 222 | Qwen 3.5 Plus | mean_score | 22 |
| 74thof 222 | Gemini 3.5 Flash-Lite | mean_score (high) | 22 |
| 75thof 222 | GLM-5.3 | mean_score (max) | 21 |
| 76thof 222 | Kimi K2.7 Code | mean_score | 21 |
| 78thof 222 | Qwen3.5-Flash | mean_score | 21 |
| 80thof 222 | Inkling | mean_score (xhigh) | 21 |
| 81stof 222 | GLM 5.2 | mean_score (max) | 21 |
| 83rdof 222 | Qwen3.6 Max Preview | mean_score (max) | 20 |
| 85thof 222 | Qwen3.6-Flash | mean_score | 20 |
| 90thof 222 | DeepSeek V4 Pro | mean_score (max) | 20 |
| 91stof 222 | GPT OSS 120B | mean_score (high) | 20 |
| 93rdof 222 | Gemini 2.5 Pro | mean_score | 20 |
| 94thof 222 | GLM 5.1 | mean_score | 19 |
| 95thof 222 | Qwen3.7 Max | mean_score (max) | 19 |
| 96thof 222 | Inkling Small | mean_score (xhigh) | 18 |
| 100thof 222 | Qwen3.6 Plus | mean_score | 17 |
| 102ndof 222 | Claude Opus 4.6 | mean_score | 17 |
| 103rdof 222 | o3-mini | mean_score (high) | 17 |
| 106thof 222 | o1 | mean_score (high) | 15 |
| 108thof 222 | GLM 5.3 Flash | mean_score (max) | 14 |
| 109thof 222 | MiniMax M3 | mean_score | 14 |
| 114thof 222 | deepseek-reasoner | mean_score | 14 |
| 115thof 222 | Seed-OSS-36B-Instruct | mean_score | 13 |
| 116thof 222 | Qwen3.5 397B A17B | mean_score (none) | 13 |
| 119thof 222 | GPT-4o (2024-08-06) | mean_score | 13 |
| 121stof 222 | Claude Sonnet 4.6 | mean_score | 13 |
| 122ndof 222 | Qwen3.5 9B | mean_score (none) | 12 |
| 123rdof 222 | Nemotron 3 Ultra | mean_score | 12 |
| 127thof 222 | GPT-5.5 Instant | mean_score | 12 |
| 128thof 222 | Kimi K2.5 | mean_score | 12 |
| 129thof 222 | Qwen3 235B A22B Thinking 2507 | mean_score | 12 |
| 130thof 222 | Claude Sonnet 4.5 | mean_score | 12 |
| 131stof 222 | Claude Opus 4.5 | mean_score | 12 |
| 132ndof 222 | Qwen3.5-35B-A3B | mean_score (none) | 10 |
| 134thof 222 | GLM-5 | mean_score | 10 |
| 140thof 222 | Qwen3 30B A3B Thinking 2507 | mean_score | 8 |
| 141stof 222 | Claude Haiku 4.5 | mean_score | 8 |
| 147thof 222 | Claude Opus 4.1 | mean_score | 7 |
| 149thof 222 | GPT-4.1 mini | mean_score | 7 |
| 151stof 222 | GPT-4.1 fine-tuned | mean_score | 6 |
| 154thof 222 | GPT-4 Turbo | mean_score | 6 |
| 155thof 222 | GLM-4.7 | mean_score | 6 |
| 156thof 222 | Qwen3 32B | mean_score | 5 |
| 157thof 222 | QwQ-32B | mean_score | 5 |
| 158thof 222 | Qwen3 8B | mean_score | 5 |
| 160thof 222 | Gemma 4 31B | mean_score (minimal) | 5 |
| 161stof 222 | Claude 3 Opus | mean_score | 5 |
| 164thof 222 | Qwen3 30B A3B | mean_score | 4 |
| 165thof 222 | Qwen3 14B | mean_score | 4 |
| 166thof 222 | Qwen3-4B-Instruct-2507 | mean_score | 4 |
| 167thof 222 | GPT OSS 20B | mean_score (medium) | 4 |
| 169thof 222 | GPT-4 (0613) | mean_score | 4 |
| 176thof 222 | DeepSeek R1 0528 Qwen3 8B | mean_score | 3 |
| 180thof 222 | Qwen3 30B A3B Instruct 2507 | mean_score | 2 |
| 186thof 222 | Mistral Small 3.1 (25.03) | mean_score | 100 |
| 188thof 222 | Qwen3 4B | mean_score (none) | 100 |
| 190thof 222 | Phi-4 | mean_score | 100 |
| 191stof 222 | DeepSeek R1 Distill QWEN 32B | mean_score | 100 |
| 192ndof 222 | DeepSeek R1 Distill QWEN 14B | mean_score | 100 |
| 193rdof 222 | deepseek-chat | mean_score | 100 |
| 195thof 222 | Granite 4.0 Micro | mean_score | 0 |
| 198thof 222 | Qwen2.5 7B Instruct | mean_score | 0 |
| 199thof 222 | AgentRL | mean_score | 0 |
| 200thof 222 | Phi-3-mini-4k-instruct | mean_score | 0 |
| 202ndof 222 | Mistral-7B-Instruct-v0.3 | mean_score | 0 |
| 203rdof 222 | Llama 3.2 1B Instruct | mean_score | 0 |
| 204thof 222 | Llama-2-7b-chat | mean_score | 0 |
| 205thof 222 | Qwen3.5-2B | mean_score (none) | 0 |
| 206thof 222 | Qwen3-1.7B | mean_score (none) | 0 |
| 207thof 222 | Llama 3 8B Instruct | mean_score | 0 |
| 208thof 222 | Llama-2-13b-chat | mean_score | 0 |
| 210thof 222 | GLM-4.7 Flash | mean_score (none) | 0 |
| 212thof 222 | Gemma 3 4B | mean_score | 0 |
| 214thof 222 | Gemma 3 12B | mean_score | 0 |
| 215thof 222 | DeepSeek-R1-Distill-Qwen-1.5B | mean_score | 0 |
| 216thof 222 | deepseek-llm-67b-chat | mean_score | 0 |
| 217thof 222 | Llama 3.1 8B Instruct | mean_score | 0 |
| 219thof 222 | Gemma 3 27B | mean_score | 0 |
| 221stof 222 | GPT-3.5 Turbo | mean_score | 0 |
| 222ndof 222 | GPT-4o-mini (2024-07-18) | mean_score | 0 |
Read from the board on 2026-09-16