Coding
Which model should write and fix code. Four answers — best value for money, best there is, best with open weights, cheapest that still works.
Which model should write and fix code. 119 measured, 113 of them for sale, and 78 of those stand on two or more of these boards — only those can be picked.
- Best value for money · Best open sourceDeepSeek-R1 (0528)DeepSeek · 5th of 28 on LiveCodeBench Leaderboard · 2 boards$0.2→$0.25per Mtok in / out13 sellers
- Best frontierClaude Fable 5Anthropic · 1st of 30 on Frontiercode · 4 boards$10→$50per Mtok in / out15 sellers
- CheapestGLM 5.3 FlashZ.ai · 10th of 121 on Webdev arena · 3 boards$0.075→$0.25per Mtok in / out17 sellers
Judged on
- SWE-Bench Pro (Public Dataset)
- SWE-bench Verified (default "Bash Only" view, agent = mini-SWE-agent)
- SWE-bench Verified (Bash Only toggle off — all agents)
- LiveCodeBench Leaderboard (default window 8/1/2024–5/1/2025, 454 problems)
- Aider polyglot coding leaderboard
- LMArena · WebDev
- Scicode — Epoch AI
- Cursorbench — Epoch AI
- Frontiercode — Epoch AI
- Terminalbench — Epoch AI
- Webdev arena — Epoch AI
- Algotune — Epoch AI
- Surface evolver bench — Epoch AI
Capability is a model's mean percentile across the boards below, on each board's widest metric. Prices are quoted as you are billed them; where a model charges separately to read and to write, the two are ranked against each other at one part in to three parts out, which for this field puts the middle at $2.19 per million tokens. Best value for money is the most capable at or under that middle, and cheapest is the least expensive of those above the middle for capability — without that second floor the cheapest answer is reliably the worst model in the category. Figures read 2026-09-16.