Evaluation models that run on 8 GB
7 models with published weights that fit in 8 GB — a phone, a base iPad, an Air. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 7 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
The memory of a phone, a base iPad or an entry-level laptop. What fits is small: models of a few billion parameters, quick and cheap to run, good at summarising, classifying and simple extraction, and out of their depth on long reasoning. This is also where on-device makes the most sense, because the alternative is a network round trip for something that takes a moment.
Tools for judging what a model produced: scoring runs, tracing a chain of calls, catching a regression before a user does. Priced per trace or per evaluation. Most use a model as the judge, which is worth knowing, because it means your evaluation has a bill and an opinion of its own.
- Qwen2.5-Omni-3BAlibaba5.5B≈3.6 GB at 4-bit
- fable-tracesAliesTaha4.0B≈2.6 GB at 4-bit
- LocoOperator-4BLocoreMind4.0B≈2.6 GB at 4-bit
- Isaac 0.2 2B PreviewPerceptron2.0B≈1.3 GB at 4-bit1 also selling it hosted
- Isaac 0.2 1BPerceptron1.0B≈0.7 GB at 4-bit1 also selling it hosted
- llama-nemotron-rerank-vl-1b-v2NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- PaddleOCR-VL-0.9Bpaddlepaddle0.9B≈0.6 GB at 4-bit1 also selling it hosted