SWE-Bench Pro (Public Dataset)
SWE-Bench Pro (Public Dataset) is run by Scale AI. It has ranked 25 models, of which the catalogue holds 25, scoring from 1.51 to 61.5 on Resolve Rate.
measured by Scale AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 25 | Muse Spark 1.1 | Resolve Rate | 61.5 |
| 1stof 25 | GPT-5.4 | Resolve Rate (xhigh) | 59.1 |
| 3rdof 25 | Muse Spark 1.2 | Resolve Rate | 55 |
| 3rdof 25 | Claude Opus 4.6 | Resolve Rate (thinking) | 51.9 |
| 5thof 25 | Gemini 3.1 Pro Preview | Resolve Rate (thinking) | 46.1 |
| 5thof 25 | Claude Opus 4.5 | Resolve Rate | 45.89 |
| 5thof 25 | Claude Sonnet 4.5 | Resolve Rate | 43.6 |
| 5thof 25 | gemini-3-pro-preview | Resolve Rate | 43.3 |
| 5thof 25 | Claude Sonnet 4 | Resolve Rate | 42.7 |
| 10thof 25 | GPT-5 | Resolve Rate (high) | 41.78 |
| 10thof 25 | GPT-5.2 Codex | Resolve Rate | 41.04 |
| 10thof 25 | Claude Haiku 4.5 | Resolve Rate | 39.45 |
| 10thof 25 | Qwen3-Coder-480B-A35B-Instruct | Resolve Rate | 38.7 |
| 14thof 25 | MiniMax M2.1 | Resolve Rate | 36.81 |
| 14thof 25 | Gemini 3 Flash Preview | Resolve Rate | 34.63 |
| 16thof 25 | GPT-5.2 | Resolve Rate | 29.94 |
| 16thof 25 | Kimi K2 0711 | Resolve Rate | 27.67 |
| 18thof 25 | Qwen3 235B A22B | Resolve Rate | 21.41 |
| 19thof 25 | GPT OSS 120B | Resolve Rate | 16.2 |
| 19thof 25 | DeepSeek v3.2 | Resolve Rate | 15.56 |
| 21stof 25 | Gemma 3 27B | Resolve Rate | 11.38 |
| 21stof 25 | Llama 3.1 405B | Resolve Rate | 11.18 |
| 21stof 25 | GLM-4.6 | Resolve Rate | 9.67 |
| 24thof 25 | Llama 4 Maverick 17B | Resolve Rate | 5.24 |
| 25thof 25 | Codestral (v24.05) | Resolve Rate | 1.51 |
Read from the board on 2026-08-25