Pass IndexThe State of AISign in

Reading documents models that run on 128 GB

24 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.

Models that read documents — scans, photographs of pages, forms, tables — and return text or structure. The interesting difference between them is not whether they can read a clean page, which they all can, but what they do with a table, a stamp, a handwritten margin or a column that breaks across pages. Several are priced per page rather than per token, which makes them easy to budget and hard to compare with the rest.

Wider

Reading documents modelsModels that run on 128 GB