Pass IndexThe State of AISign in

SWE-bench Verified (default "Bash Only" view, agent = mini-SWE-agent)

SWE-bench Verified (default "Bash Only" view, agent = mini-SWE-agent) is run by SWE-bench. It has ranked 47 models, of which the catalogue holds 24, scoring from 60 to 76.8 on % Resolved.

measured by SWE-bench · the board itself

PlaceModelMetricScore
1stof 47Claude Opus 4.5% Resolved (high)76.8
2ndof 47Gemini 3 Flash Preview% Resolved (high)75.8
3rdof 47MiniMax M2.5% Resolved (high)75.8
4thof 47Claude Opus 4.6% Resolved75.6
5thof 47Claude Opus 4.5% Resolved (medium)74.4
6thof 47gemini-3-pro-preview% Resolved74.2
7thof 47GPT-5.2 Codex% Resolved72.8
8thof 47GLM-5% Resolved (high)72.8
9thof 47GPT-5.2% Resolved72.8
9thof 47GPT-5.2% Resolved (high)72.8
11thof 47Claude Sonnet 4.5% Resolved (high)71.4
11thof 47Claude Sonnet 4.5% Resolved71.4
12thof 47Kimi K2.5% Resolved (high)70.8
14thof 47DeepSeek v3.2% Resolved (high)70
15thof 47gemini-3-pro-preview% Resolved (high)69.6
17thof 47Claude Opus 4% Resolved67.6
18thof 47Claude Haiku 4.5% Resolved (high)66.6
19thof 47GPT-5.1 Codex% Resolved (medium)66
20thof 47GPT-5.1% Resolved (medium)66
21stof 47GPT-5% Resolved (medium)65
22ndof 47Claude Sonnet 4% Resolved64.93
23rdof 47Kimi K2 Thinking% Resolved63.4
24thof 47MiniMax-M2% Resolved61
25thof 47DeepSeek v3.2% Resolved (reasoner)60

Read from the board on 2026-08-25