Deepresearchbench — Epoch AI
Deepresearchbench — Epoch AI is run by Epoch AI. It has ranked 41 models, of which the catalogue holds 24, scoring from 0.351 to 0.553.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 41 | Claude Opus 4.6 | Average score (high) | 0.553 |
| 2ndof 41 | Claude Sonnet 4.6 | Average score (high) | 0.549 |
| 3rdof 41 | Claude Opus 4.5 | Average score (high) | 0.548 |
| 4thof 41 | GPT-5.5 | Average score (high) | 0.54 |
| 7thof 41 | Claude Sonnet 4.5 | Average score | 0.526 |
| 10thof 41 | Claude Opus 4.8 | Average score (high) | 0.502 |
| 11thof 41 | Gemini 3 Flash | Average score (low) | 0.498 |
| 12thof 41 | GPT-5 | Average score (low) | 0.496 |
| 19thof 41 | Claude Opus 4.1 | Average score | 0.483 |
| 22ndof 41 | Gemini 3.1 Pro Preview | Average score (high) | 0.478 |
| 26thof 41 | Claude Opus 4 | Average score | 0.468 |
| 27thof 41 | Claude Sonnet 4 | Average score | 0.466 |
| 28thof 41 | gemini-3-pro-preview | Average score (low) | 0.463 |
| 29thof 41 | Claude Haiku 4.5 | Average score (low) | 0.455 |
| 30thof 41 | o3 | Average score (medium) | 0.452 |
| 32ndof 41 | Claude 3.7 Sonnet | Average score | 0.436 |
| 33rdof 41 | Gemini 2.5 Pro Preview 06-05 | Average score | 0.428 |
| 34thof 41 | GPT-5.1 | Average score (low) | 0.428 |
| 35thof 41 | Gemini 2.5 Pro | Average score | 0.415 |
| 36thof 41 | GPT-5.2 | Average score (low) | 0.411 |
| 37thof 41 | Gemini 3.1 Flash-Lite | Average score (low) | 0.373 |
| 39thof 41 | GPT-5.4 mini | Average score (low) | 0.363 |
| 40thof 41 | GPT-5.4 | Average score (low) | 0.351 |
| 41stof 41 | DeepSeek-R1 (0528) | Average score | 0.351 |
Read from the board on 2026-09-16