Open weights
Which model to run yourself, or buy from whoever you like. Four answers — best value for money, best there is, best with open weights, cheapest that still works.
Which model to run yourself, or buy from whoever you like. 203 measured, 132 of them for sale, and 89 of those stand on two or more of these boards — only those can be picked.
Judged on
- Artificial Analysis · Intelligence Index
- AgentBench Leaderboard (new)
- Aider polyglot coding leaderboard
- AIME 2025
- AIME 2026
- ARC-AGI-1 (Semi-Private)
- ARC-AGI-1 (Public Eval)
- ARC-AGI-2 (Semi-Private)
- ARC-AGI-2 (Public Eval)
- Berkeley Function-Calling Leaderboard (BFCL) V4
- Algotune — Epoch AI
- Apex agents — Epoch AI
- Arc agi — Epoch AI
- Arc agi 2 — Epoch AI
- Arc ai2 — Epoch AI
- Blueprint bench 2 — Epoch AI
- Epoch capabilities index — Epoch AI
- Chess puzzles — Epoch AI
- Critpt — Epoch AI
- Cursorbench — Epoch AI
- Enigma eval — Epoch AI
- Forecastbench — Epoch AI
- Frontiercode — Epoch AI
- FrontierMath (Tiers 1-3)
- Frontiermath tier 4 — Epoch AI
- Gbaeval — Epoch AI
- Humanity's Last Exam (Epoch AI replication)
- Lambada — Epoch AI
- Lech mazur writing — Epoch AI
- Math level 5 — Epoch AI
- Metr time horizons — Epoch AI
- Mystery game puzzles — Epoch AI
- Piqa — Epoch AI
- Proofbench — Epoch AI
- Rli — Epoch AI
- Scicode — Epoch AI
- Science qa — Epoch AI
- Simplebench — Epoch AI
- Simpleqa verified — Epoch AI
- Surface evolver bench — Epoch AI
- Terminalbench — Epoch AI
- The agent company — Epoch AI
- Vending bench 2 — Epoch AI
- Webdev arena — Epoch AI
- Weirdml — Epoch AI
- Wino grande — Epoch AI
- GPQA Diamond
- Humanity's Last Exam
- LiveCodeBench Leaderboard (default window 8/1/2024–5/1/2025, 454 problems)
- LMArena · Agent
- LMArena · Search
- LMArena · Text
- LMArena · Vision
- LMArena · WebDev
- MTEB · English (MTEB(eng, v2))
- MTEB · English — Reranking task-type split
- MTEB · Multilingual (MTEB(Multilingual, v2))
- OSWorld-Verified Results (All)
- SWE-Bench Pro (Public Dataset)
- SWE-bench Verified (default "Bash Only" view, agent = mini-SWE-agent)
- SWE-bench Verified (Bash Only toggle off — all agents)
- τ³-Banking Leaderboard
- τ³-Voice Leaderboard
- Vending-Bench 2 — Current leaderboard
- WebArena Leaderboard (official Google Sheet)
Capability is a model's mean percentile across the boards below, on each board's widest metric. Prices are quoted as you are billed them; where a model charges separately to read and to write, the two are ranked against each other at one part in to three parts out, which for this field puts the middle at $0.4 per million tokens. Best value for money is the most capable at or under that middle, and cheapest is the least expensive of those above the middle for capability — without that second floor the cheapest answer is reliably the worst model in the category. Figures read 2026-09-16.