Swe bench verified — Epoch AI
Swe bench verified — Epoch AI is run by Epoch AI. It has ranked 35 models, of which the catalogue holds 32, scoring from 30.992 to 83.471.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 35 | Claude Opus 4.7 | mean_score (max) | 83.471 |
| 2ndof 35 | GPT-5.5 | mean_score (xhigh) | 80.579 |
| 3rdof 35 | Gemini 3.5 Flash | mean_score (high) | 79.339 |
| 4thof 35 | Claude Opus 4.6 | mean_score | 78.719 |
| 5thof 35 | GLM 5.2 | mean_score (max) | 78.7 |
| 6thof 35 | DeepSeek V4 Pro | mean_score (max) | 77.64 |
| 7thof 35 | Qwen3.7 Max | mean_score (max) | 77.273 |
| 8thof 35 | GPT-5.4 | mean_score (high) | 76.86 |
| 9thof 35 | Qwen3.6 Max Preview | mean_score (max) | 76.653 |
| 10thof 35 | Kimi K2.6 | mean_score | 76.653 |
| 11thof 35 | Claude Opus 4.5 | mean_score | 76.653 |
| 12thof 35 | Gemini 3.1 Pro Preview Custom Tools | mean_score | 75.62 |
| 14thof 35 | Gemini 3 Flash Preview | mean_score | 75.413 |
| 15thof 35 | Claude Sonnet 4.6 | mean_score | 75.207 |
| 16thof 35 | GPT-5.3 Codex | mean_score (high) | 74.793 |
| 17thof 35 | GLM 5.1 | mean_score | 74.17 |
| 18thof 35 | Kimi K2.5 | mean_score | 73.76 |
| 19thof 35 | GPT-5.2 | mean_score (high) | 73.76 |
| 20thof 35 | GPT-5 | mean_score (high) | 73.554 |
| 21stof 35 | Claude Opus 4.1 | mean_score | 73.347 |
| 22ndof 35 | gemini-3-pro-preview | mean_score | 72.934 |
| 23rdof 35 | GLM-5 | mean_score | 72.078 |
| 25thof 35 | Claude Sonnet 4.5 | mean_score | 71.281 |
| 26thof 35 | Claude Opus 4 | mean_score | 70.661 |
| 27thof 35 | GPT-5.1 | mean_score (high) | 67.975 |
| 29thof 35 | GPT-5 mini | mean_score (medium) | 64.669 |
| 30thof 35 | o3 | mean_score (medium) | 62.319 |
| 31stof 35 | Claude 3.7 Sonnet | mean_score | 60.95 |
| 32ndof 35 | Qwen3.6 Plus | mean_score | 57.851 |
| 33rdof 35 | Gemini 2.5 Pro | mean_score | 57.557 |
| 34thof 35 | GPT-4.1 fine-tuned | mean_score | 48.542 |
| 35thof 35 | GPT-4o (2024-11-20) | mean_score | 30.992 |
Read from the board on 2026-09-16