Pass IndexThe State of AISign in

SWE-Bench Pro (Public Dataset)

SWE-Bench Pro (Public Dataset) is run by Scale AI. It has ranked 25 models, of which the catalogue holds 25, scoring from 1.51 to 61.5 on Resolve Rate.

measured by Scale AI · the board itself

PlaceModelMetricScore
1stof 25Muse Spark 1.1Resolve Rate61.5
1stof 25GPT-5.4Resolve Rate (xhigh)59.1
3rdof 25Muse Spark 1.2Resolve Rate55
3rdof 25Claude Opus 4.6Resolve Rate (thinking)51.9
5thof 25Gemini 3.1 Pro PreviewResolve Rate (thinking)46.1
5thof 25Claude Opus 4.5Resolve Rate45.89
5thof 25Claude Sonnet 4.5Resolve Rate43.6
5thof 25gemini-3-pro-previewResolve Rate43.3
5thof 25Claude Sonnet 4Resolve Rate42.7
10thof 25GPT-5Resolve Rate (high)41.78
10thof 25GPT-5.2 CodexResolve Rate41.04
10thof 25Claude Haiku 4.5Resolve Rate39.45
10thof 25Qwen3-Coder-480B-A35B-InstructResolve Rate38.7
14thof 25MiniMax M2.1Resolve Rate36.81
14thof 25Gemini 3 Flash PreviewResolve Rate34.63
16thof 25GPT-5.2Resolve Rate29.94
16thof 25Kimi K2 0711Resolve Rate27.67
18thof 25Qwen3 235B A22BResolve Rate21.41
19thof 25GPT OSS 120BResolve Rate16.2
19thof 25DeepSeek v3.2Resolve Rate15.56
21stof 25Gemma 3 27BResolve Rate11.38
21stof 25Llama 3.1 405BResolve Rate11.18
21stof 25GLM-4.6Resolve Rate9.67
24thof 25Llama 4 Maverick 17BResolve Rate5.24
25thof 25Codestral (v24.05)Resolve Rate1.51

Read from the board on 2026-08-25