Reasoning
Which model can do the mathematics and the science. Four answers — best value for money, best there is, best with open weights, cheapest that still works.
Which model can do the mathematics and the science. 197 measured, 188 of them for sale, and 164 of those stand on two or more of these boards — only those can be picked.
- Best value for money · Best open sourceStep 3.5 FlashStepFun · 3rd of 61 on AIME 2025 · 2 boards$0.09→$0.3per Mtok in / out8 sellers
- Best frontierClaude Opus 5Anthropic · 1st of 23 on Gbaeval · 3 boards$5→$25per Mtok in / out15 sellers
- CheapestDeepSeek V4 Flash (0731)DeepSeek · 18th of 127 on Mystery game puzzles · 3 boards$0.04→$0.1per Mtok in / out18 sellers
Judged on
- FrontierMath (Tiers 1-3)
- Humanity's Last Exam (Epoch AI replication)
- Humanity's Last Exam
- GPQA Diamond
- AIME 2026
- AIME 2025
- ARC-AGI-1 (Semi-Private)
- ARC-AGI-2 (Semi-Private)
- Artificial Analysis · Intelligence Index
- Epoch capabilities index — Epoch AI
- Weirdml — Epoch AI
- Critpt — Epoch AI
- Chess puzzles — Epoch AI
- Math level 5 — Epoch AI
- Arc agi — Epoch AI
- Arc agi 2 — Epoch AI
- Simplebench — Epoch AI
- Proofbench — Epoch AI
- Frontiermath tier 4 — Epoch AI
- Mystery game puzzles — Epoch AI
- Enigma eval — Epoch AI
- Gbaeval — Epoch AI
- ARC-AGI-1 (Public Eval)
- ARC-AGI-2 (Public Eval)
Capability is a model's mean percentile across the boards below, on each board's widest metric. Prices are quoted as you are billed them; where a model charges separately to read and to write, the two are ranked against each other at one part in to three parts out, which for this field puts the middle at $1.04 per million tokens. Best value for money is the most capable at or under that middle, and cheapest is the least expensive of those above the middle for capability — without that second floor the cheapest answer is reliably the worst model in the category. Figures read 2026-09-16.