Pass IndexThe State of AISign in

Reading documents models that run on 24 GB

24 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.

Models that read documents — scans, photographs of pages, forms, tables — and return text or structure. The interesting difference between them is not whether they can read a clean page, which they all can, but what they do with a table, a stamp, a handwritten margin or a column that breaks across pages. Several are priced per page rather than per token, which makes them easy to budget and hard to compare with the rest.

Wider

Reading documents modelsModels that run on 24 GB