Pass IndexThe State of AISign in

Reasoning models that run on 64 GB

62 models with published weights that fit in 64 GB — a workstation. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 67 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

A workstation, or a well-specified Mac. Seventy-billion-parameter models fit at four-bit, which is the class where local output stops being obviously worse than the hosted models people pay for. Loading takes real time and the machine will be warm.

Models that plan before they answer, spending extra tokens on working the problem through. They win on mathematics, hard code and anything with several steps, and they are measured on boards like GPQA, AIME and ARC-AGI. The catch is what the thinking costs: reasoning is billed as output tokens, so the same model can cost several times more per answer at a high effort than a low one, and for a question that needed one turn you have paid for a monologue.

Wider

Reasoning modelsModels that run on 64 GB