Terminalbench — Epoch AI
Terminalbench — Epoch AI is run by Epoch AI. It has ranked 202 models, of which the catalogue holds 43, scoring from 0.034 to 0.847.
measured by Epoch AI · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 202 | GPT-5.5 | Accuracy mean | 0.847 |
| 5thof 202 | GPT-5.4 | Accuracy mean | 0.818 |
| 6thof 202 | Gemini 3.1 Pro Preview | Accuracy mean | 0.802 |
| 7thof 202 | Claude Opus 4.7 | Accuracy mean | 0.802 |
| 8thof 202 | Claude Opus 4.6 | Accuracy mean | 0.798 |
| 11thof 202 | GPT-5.3 Codex | Accuracy mean | 0.784 |
| 27thof 202 | gemini-3-pro-preview | Accuracy mean | 0.694 |
| 31stof 202 | GPT-5.2 Codex | Accuracy mean | 0.665 |
| 34thof 202 | GPT-5.2 | Accuracy mean (medium) | 0.649 |
| 37thof 202 | Gemini 3 Flash Preview | Accuracy mean | 0.643 |
| 38thof 202 | Claude Opus 4.5 | Accuracy mean | 0.631 |
| 46thof 202 | GPT-5.1 Codex Mini | Accuracy mean | 0.616 |
| 49thof 202 | GPT-5.1 Codex Max | Accuracy mean (max) | 0.604 |
| 55thof 202 | GPT-5.1 Codex | Accuracy mean | 0.578 |
| 58thof 202 | Grok 4.20 | Accuracy mean | 0.573 |
| 65thof 202 | Claude Sonnet 4.6 | Accuracy mean | 0.534 |
| 67thof 202 | GLM-5 | Accuracy mean | 0.524 |
| 73rdof 202 | GPT-5 | Accuracy mean | 0.496 |
| 75thof 202 | GPT-5.1 | Accuracy mean | 0.476 |
| 79thof 202 | Claude Sonnet 4.5 | Accuracy mean | 0.465 |
| 80thof 202 | MiniMax M2.7 | Accuracy mean | 0.451 |
| 81stof 202 | GPT-5 Codex (batch) | Accuracy mean | 0.443 |
| 87thof 202 | Kimi K2.5 | Accuracy mean | 0.432 |
| 95thof 202 | MiniMax M2.5 | Accuracy mean | 0.427 |
| 106thof 202 | DeepSeek v3.2 | Accuracy mean | 0.396 |
| 107thof 202 | Claude Opus 4.1 | Accuracy mean | 0.38 |
| 113thof 202 | MiniMax M2.1 | Accuracy mean | 0.366 |
| 114thof 202 | Kimi K2 Thinking | Accuracy mean | 0.357 |
| 116thof 202 | Claude Haiku 4.5 | Accuracy mean | 0.355 |
| 123rdof 202 | GPT-5 mini | Accuracy mean | 0.348 |
| 127thof 202 | GLM-4.7 | Accuracy mean | 0.334 |
| 129thof 202 | Gemini 2.5 Pro | Accuracy mean | 0.326 |
| 133rdof 202 | MiniMax-M2 | Accuracy mean | 0.3 |
| 141stof 202 | Kimi K2 Instruct | Accuracy mean | 0.278 |
| 147thof 202 | Qwen3-Coder-480B-A35B-Instruct | Accuracy mean | 0.272 |
| 152ndof 202 | grok-code-fast-1 | Accuracy mean | 0.258 |
| 158thof 202 | GLM-4.6 | Accuracy mean | 0.245 |
| 166thof 202 | Qwen3.6 35B A3B | Accuracy mean | 0.23 |
| 169thof 202 | GPT-5 nano | Accuracy mean | 0.218 |
| 172ndof 202 | GPT OSS 120B | Accuracy mean | 0.187 |
| 175thof 202 | Gemini 2.5 Flash | Accuracy mean | 0.171 |
| 194thof 202 | Qwen3.5 9B | Accuracy mean | 0.092 |
| 199thof 202 | GPT OSS 20B | Accuracy mean | 0.034 |
Read from the board on 2026-09-16