Structured extraction models that run on 16 GB
15 models with published weights that fit in 16 GB — the common laptop. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 16 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
The commonest laptop configuration, and the point where a genuinely useful model runs locally: fourteen billion parameters at four-bit fits with room to spare for the context. Expect a capable general assistant that will not match the frontier, and remember that anything else the machine is doing competes for the same memory.
Tools that pull structure out of unstructured input — fields from an invoice, a schema from a page, a table from a report. Some are models, some are services with a model inside. Judge them on what they do when the input is malformed, because that is the whole job; anything can parse a clean file.
- Gemma-3-R1984-12BVIDraft12.2B≈7.9 GB at 4-bit
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Qwen3.5 9BAlibaba9.7B≈6.3 GB at 4-bit8 also selling it hosted
- liftDatalab9.7B≈6.3 GB at 4-bit
- MiniCPM-SALAOpenBMB9.5B≈6.2 GB at 4-bit
- AutoGLM-Phone-9B-MultilingualZ.ai9.0B≈5.9 GB at 4-bit3 also selling it hosted
- Llama-3-RefueledRefuel AI8.0B≈5.2 GB at 4-bit
- NuExtract3NuMind4.5B≈3.0 GB at 4-bit
- NuExtractNuMind3.8B≈2.5 GB at 4-bit
- NuExtract-1.5NuMind3.8B≈2.5 GB at 4-bit
- dots.ocrdots studio3.0B≈2.0 GB at 4-bit
- DeepScaleR-1.5B-PreviewAgentica1.8B≈1.2 GB at 4-bit
- MinerU2.5-Pro-2604-1.2BOpenDataLab1.2B≈0.8 GB at 4-bit
- Qwen3.5-0.8BAlibaba0.8B≈0.5 GB at 4-bit1 also selling it hosted
- DolphinByteDance0.4B≈0.3 GB at 4-bit