Structured extraction models that run on 8 GB
8 models with published weights that fit in 8 GB — a phone, a base iPad, an Air. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 7 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
The memory of a phone, a base iPad or an entry-level laptop. What fits is small: models of a few billion parameters, quick and cheap to run, good at summarising, classifying and simple extraction, and out of their depth on long reasoning. This is also where on-device makes the most sense, because the alternative is a network round trip for something that takes a moment.
Tools that pull structure out of unstructured input — fields from an invoice, a schema from a page, a table from a report. Some are models, some are services with a model inside. Judge them on what they do when the input is malformed, because that is the whole job; anything can parse a clean file.
- NuExtract3NuMind4.5B≈3.0 GB at 4-bit
- NuExtractNuMind3.8B≈2.5 GB at 4-bit
- NuExtract-1.5NuMind3.8B≈2.5 GB at 4-bit
- dots.ocrdots studio3.0B≈2.0 GB at 4-bit
- DeepScaleR-1.5B-PreviewAgentica1.8B≈1.2 GB at 4-bit
- MinerU2.5-Pro-2604-1.2BOpenDataLab1.2B≈0.8 GB at 4-bit
- Qwen3.5-0.8BAlibaba0.8B≈0.5 GB at 4-bit1 also selling it hosted
- DolphinByteDance0.4B≈0.3 GB at 4-bit