Bang for the buck
The three best models in each of the four categories — best by that category's own leaderboards, benchmarks first — among those costing under a dollar the million tokens. And the average of their sizes: how big a model leads its category at that price.
Overall
298B
average size across the four tops
$0.5
average price, the million tokens out
12 models picked across the four categories
General
11B
average size of the top three by benchmarks
- GPT-4o-mini (2024-07-18)about 14B$0.4
- Phi-3-mini-4k-instruct3.8B$0.52
- DeepSeek-V3about 14B$0.2
Reasoning
219B
average size of the top three by benchmarks
- MiniMax M2.1229B$0.95
- Seed-2.0-Liteabout 229B$0.54
- Step 3.5 Flash199B$0.3
Coding
279B
average size of the top three by benchmarks
- MiniMax M2.5229B$0.9
- GLM 5.3 Flashabout 304B$0.25
- DeepSeek V4.1 Flashabout 304B$0.31
Agentic
685B
average size of the top three by benchmarks
- DeepSeek v3.2685B$0.31
- AgentRLabout 685B$0.8
- Grok 4.1 Fast Reasoningabout 685B$0.5
How this is worked out
A model counts if a company sells its output for under a dollar the million tokens on the ordinary lane — the current rate, not the cheapest ever recorded — and if it places on one of that category's own boards. The three shown are the top three by benchmarks: sorted by the model's best placing among that category's boards, as a share of the field, price breaking ties.
Most leaders at this price are closed and publish no parameter count. Leaving them out would answer a different question, so each is shown as about the median size of the three models nearest it in standing that do publish one. Every such figure is marked; the rest are the makers' own.
every model under a dollar · the categories · models by size · this page as JSON