Pass IndexThe State of AISign in

Evaluation models that run on 8 GB

7 models with published weights that fit in 8 GB — a phone, a base iPad, an Air. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 7 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

The memory of a phone, a base iPad or an entry-level laptop. What fits is small: models of a few billion parameters, quick and cheap to run, good at summarising, classifying and simple extraction, and out of their depth on long reasoning. This is also where on-device makes the most sense, because the alternative is a network round trip for something that takes a moment.

Tools for judging what a model produced: scoring runs, tracing a chain of calls, catching a regression before a user does. Priced per trace or per evaluation. Most use a model as the judge, which is worth knowing, because it means your evaluation has a bill and an opinion of its own.

Wider

Evaluation modelsModels that run on 8 GB