SWE-bench Verified (default "Bash Only" view, agent = mini-SWE-agent)
SWE-bench Verified (default "Bash Only" view, agent = mini-SWE-agent) is run by SWE-bench. It has ranked 47 models, of which the catalogue holds 24, scoring from 60 to 76.8 on % Resolved.
measured by SWE-bench · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 47 | Claude Opus 4.5 | % Resolved (high) | 76.8 |
| 2ndof 47 | Gemini 3 Flash Preview | % Resolved (high) | 75.8 |
| 3rdof 47 | MiniMax M2.5 | % Resolved (high) | 75.8 |
| 4thof 47 | Claude Opus 4.6 | % Resolved | 75.6 |
| 5thof 47 | Claude Opus 4.5 | % Resolved (medium) | 74.4 |
| 6thof 47 | gemini-3-pro-preview | % Resolved | 74.2 |
| 7thof 47 | GPT-5.2 Codex | % Resolved | 72.8 |
| 8thof 47 | GLM-5 | % Resolved (high) | 72.8 |
| 9thof 47 | GPT-5.2 | % Resolved | 72.8 |
| 9thof 47 | GPT-5.2 | % Resolved (high) | 72.8 |
| 11thof 47 | Claude Sonnet 4.5 | % Resolved (high) | 71.4 |
| 11thof 47 | Claude Sonnet 4.5 | % Resolved | 71.4 |
| 12thof 47 | Kimi K2.5 | % Resolved (high) | 70.8 |
| 14thof 47 | DeepSeek v3.2 | % Resolved (high) | 70 |
| 15thof 47 | gemini-3-pro-preview | % Resolved (high) | 69.6 |
| 17thof 47 | Claude Opus 4 | % Resolved | 67.6 |
| 18thof 47 | Claude Haiku 4.5 | % Resolved (high) | 66.6 |
| 19thof 47 | GPT-5.1 Codex | % Resolved (medium) | 66 |
| 20thof 47 | GPT-5.1 | % Resolved (medium) | 66 |
| 21stof 47 | GPT-5 | % Resolved (medium) | 65 |
| 22ndof 47 | Claude Sonnet 4 | % Resolved | 64.93 |
| 23rdof 47 | Kimi K2 Thinking | % Resolved | 63.4 |
| 24thof 47 | MiniMax-M2 | % Resolved | 61 |
| 25thof 47 | DeepSeek v3.2 | % Resolved (reasoner) | 60 |
Read from the board on 2026-08-25