Pass IndexThe State of AISign in

Conversation models that run on 32 GB

515 models with published weights that fit in 32 GB — a well-specified laptop. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 33 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 515 in all, a hundred to a page; this is page 5 of 6.

Enough for a thirty-billion-parameter model at four-bit with a long context, or a smaller one at higher precision if quality matters more than size. A practical ceiling for a laptop that also has to be a laptop.

Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.

Wider

Conversation modelsModels that run on 32 GB