AgentBench Leaderboard (new)
AgentBench Leaderboard (new) is run by THUDM / Tsinghua KEG. It has ranked 25 models, of which the catalogue holds 18, scoring from 36.1 to 70.4 on Success Rate (pass@1) AVG.
measured by THUDM / Tsinghua KEG · the board itself
| Place | Model | Metric | Score |
|---|---|---|---|
| 1stof 25 | AgentRL | Success Rate (pass@1) AVG | 70.4 |
| 4thof 25 | Qwen 2.5 7B Instruct Turbo | Success Rate (pass@1) AVG | 62 |
| 6thof 25 | Claude Sonnet 4.5 | Success Rate (pass@1) AVG | 58.9 |
| 7thof 25 | Claude Sonnet 4.5 | Success Rate (pass@1) AVG (thinking) | 58.3 |
| 8thof 25 | Claude Sonnet 4 | Success Rate (pass@1) AVG | 58.2 |
| 8thof 25 | Claude Sonnet 4 | Success Rate (pass@1) AVG (thinking) | 58.2 |
| 10thof 25 | Claude 3.7 Sonnet | Success Rate (pass@1) AVG | 53.2 |
| 11thof 25 | GPT-5 | Success Rate (pass@1) AVG | 52.2 |
| 12thof 25 | AgentLM-70B | Success Rate (pass@1) AVG | 51.4 |
| 13thof 25 | Claude 3.7 Sonnet | Success Rate (pass@1) AVG (thinking) | 50 |
| 14thof 25 | DeepSeek-R1 | Success Rate (pass@1) AVG | 49.3 |
| 15thof 25 | AgentLM-13B | Success Rate (pass@1) AVG | 45.1 |
| 16thof 25 | AgentLM-7B | Success Rate (pass@1) AVG | 42.7 |
| 17thof 25 | o3-mini | Success Rate (pass@1) AVG | 40.9 |
| 18thof 25 | Qwen2.5 72B Instruct | Success Rate (pass@1) AVG | 40.8 |
| 19thof 25 | o4-mini | Success Rate (pass@1) AVG | 39.7 |
| 20thof 25 | GPT-4o (2024-11-20) | Success Rate (pass@1) AVG | 39.6 |
| 23rdof 25 | DeepSeek-V3 | Success Rate (pass@1) AVG | 36.1 |
Read from the board on 2026-08-25