Agents
Which model can call tools and finish a job. Four answers — best value for money, best there is, best with open weights, cheapest that still works.
Which model can call tools and finish a job. 152 measured, 135 of them for sale, and 76 of those stand on two or more of these boards — only those can be picked.
- Best value for moneyMuse Spark 1.1Meta · 7th of 110 on OSWorld-Verified Results · 5 boards$1.25→$4.25per Mtok in / out4 sellers
- Best frontierGPT 6 AstraOpenAI · 2nd of 65 on Apex agents · 2 boards$10→$50per Mtok in / out15 sellers
- Best open sourceGLM 5.2Z.ai · 5th of 60 on Vending-Bench 2 · 4 boards$0.49→$1.73per Mtok in / out24 sellers
- CheapestDeepSeek V3.2 ExpDeepSeek · 1st of 16 on The agent company · 3 boards$0.21→$0.32per Mtok in / out8 sellers
Judged on
- Berkeley Function-Calling Leaderboard (BFCL) V4
- τ³-Banking Leaderboard
- AgentBench Leaderboard (new)
- LMArena · Agent
- Vending-Bench 2 — Current leaderboard
- OSWorld-Verified Results (All)
- WebArena Leaderboard (official Google Sheet)
- Apex agents — Epoch AI
- The agent company — Epoch AI
- Blueprint bench 2 — Epoch AI
- Forecastbench — Epoch AI
- Metr time horizons — Epoch AI
- Rli — Epoch AI
- Vending bench 2 — Epoch AI
Capability is a model's mean percentile across the boards below, on each board's widest metric. Prices are quoted as you are billed them; where a model charges separately to read and to write, the two are ranked against each other at one part in to three parts out, which for this field puts the middle at $3.58 per million tokens. Best value for money is the most capable at or under that middle, and cheapest is the least expensive of those above the middle for capability — without that second floor the cheapest answer is reliably the worst model in the category. Figures read 2026-09-16.