Pass IndexThe State of AISign in

Conversation models that run on 16 GB

434 models with published weights that fit in 16 GB — the common laptop. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 16 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 434 in all, a hundred to a page; this is page 3 of 5.

The commonest laptop configuration, and the point where a genuinely useful model runs locally: fourteen billion parameters at four-bit fits with room to spare for the context. Expect a capable general assistant that will not match the frontier, and remember that anything else the machine is doing competes for the same memory.

Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.

Wider

Conversation modelsModels that run on 16 GB