ARC-AGI-1 (Public Eval)
ARC-AGI-1 (Public Eval) is run by ARC Prize Foundation. It has ranked 205 models, of which the catalogue holds 58, scoring from 1.75 to 99.
measured by ARC Prize Foundation · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 205 | GPT-5.6 Sol | Score (%) (max) | 99 |
| 2ndof 205 | Claude Opus 5 | Score (%) (high) | 99 |
| 4thof 205 | Claude Fable 5.1 | Score (%) (max) | 99 |
| 5thof 205 | GPT 6 Astra | Score (%) (xhigh) | 99 |
| 10thof 205 | GPT-5.4 Pro | Score (%) (xhigh) | 98.25 |
| 11thof 205 | GPT-5.6 Terra | Score (%) (max) | 98.25 |
| 13thof 205 | GPT-5.5 Pro | Score (%) (high) | 98 |
| 14thof 205 | Claude Fable 5 | Score (%) (high) | 98 |
| 21stof 205 | GPT-5.2 Pro | Score (%) (xhigh) | 97.61 |
| 25thof 205 | Gemini 3.1 Pro Preview | Score (%) | 97.24 |
| 26thof 205 | Claude Opus 4.7 | Score (%) (max) | 97 |
| 27thof 205 | Claude Opus 4.6 | Score (%) (max) | 96.75 |
| 29thof 205 | Gemini 3.7 Flash | Score (%) (high) | 96.62 |
| 33rdof 205 | GPT-5.4 | Score (%) (xhigh) | 96.38 |
| 35thof 205 | Grok 4.6 | Score (%) (xhigh) | 96.25 |
| 38thof 205 | Gemini 3.5 Flash | Score (%) (high) | 96 |
| 39thof 205 | Gemini 3.6 Flash | Score (%) (high) | 96 |
| 40thof 205 | Claude Sonnet 4.6 | Score (%) (max) | 95.75 |
| 45thof 205 | DeepSeek V4 Pro 0813 | Score (%) (low) | 95.5 |
| 50thof 205 | Kimi K3 | Score (%) (max) | 94.88 |
| 52ndof 205 | Grok 4.5 | Score (%) (high) | 94.75 |
| 53rdof 205 | DeepSeek V4 Flash (0731) | Score (%) (max) | 94.75 |
| 67thof 205 | GPT-5.6 Luna | Score (%) (max) | 91.25 |
| 70thof 205 | Inkling | Score (%) | 90.62 |
| 75thof 205 | Inkling Small | Score (%) (xhigh) | 89.38 |
| 89thof 205 | GLM 5.2 | Score (%) | 80.38 |
| 93rdof 205 | GPT 5.1 Thinking | Score (%) (high) | 77.12 |
| 94thof 205 | GPT-5 Pro | Score (%) | 77 |
| 96thof 205 | GPT-5.4 mini | Score (%) (xhigh) | 75.12 |
| 99thof 205 | Kimi K2.5 | Score (%) | 73.12 |
| 103rdof 205 | o4-mini | Score (%) (high) | 68.03 |
| 105thof 205 | Gemini 3.5 Flash-Lite | Score (%) (high) | 66.38 |
| 109thof 205 | GPT-5 | Score (%) (high) | 65.88 |
| 111thof 205 | o3 | Score (%) (high) | 64.25 |
| 115thof 205 | o3-pro | Score (%) (high) | 63.34 |
| 117thof 205 | DeepSeek v3.2 | Score (%) | 61.62 |
| 119thof 205 | GPT-5 mini | Score (%) (high) | 61.52 |
| 120thof 205 | MiniMax M2.5 | Score (%) | 59.13 |
| 121stof 205 | GLM-5 | Score (%) | 58.63 |
| 125thof 205 | Claude Sonnet 4 | Score (%) | 56.75 |
| 130thof 205 | Claude Opus 4 | Score (%) | 54.25 |
| 133rdof 205 | GPT-5.4 nano | Score (%) (high) | 51.62 |
| 146thof 205 | o3-mini | Score (%) (high) | 46.58 |
| 156thof 205 | Gemini 2.5 Pro | Score (%) | 43 |
| 160thof 205 | Gemini 2.5 Flash | Score (%) | 36.3 |
| 163rdof 205 | Claude Sonnet 4.5 | Score (%) | 35.38 |
| 165thof 205 | Codex mini | Score (%) | 33.38 |
| 171stof 205 | GPT-5 nano | Score (%) (high) | 29.67 |
| 174thof 205 | DeepSeek-R1 (0528) | Score (%) | 26.98 |
| 175thof 205 | Claude Haiku 4.5 | Score (%) | 26.62 |
| 183rdof 205 | Grok 3 Mini | Score (%) (low) | 17.62 |
| 186thof 205 | Qwen3-235B-A22B-Instruct-2507 | Score (%) | 17 |
| 190thof 205 | o1-mini | Score (%) | 13.16 |
| 193rdof 205 | GPT-4.1 fine-tuned | Score (%) | 11.75 |
| 197thof 205 | Magistral Small 2506 | Score (%) | 8.62 |
| 198thof 205 | Grok 3 | Score (%) | 8.38 |
| 200thof 205 | GPT-4.1 mini | Score (%) | 7.25 |
| 205thof 205 | GPT-4.1 nano | Score (%) | 1.75 |
Read from the board on 2026-09-16