Pass IndexThe State of AISign in

Conversation models that run on 36 GB

539 models with published weights that fit in 36 GB — a MacBook Pro with 36. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 37 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 539 in all, a hundred to a page; this is page 3 of 6.

An Apple configuration, and a comfortable one: it holds what 32 GB holds without the machine feeling tight, which in practice means a longer context or a browser you do not have to close first.

Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.

Wider

Conversation modelsModels that run on 36 GB