Structured extraction models that run on 24 GB
15 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
Tools that pull structure out of unstructured input — fields from an invoice, a schema from a page, a table from a report. Some are models, some are services with a model inside. Judge them on what they do when the input is malformed, because that is the whole job; anything can parse a clean file.
- Gemma-3-R1984-12BVIDraft12.2B≈7.9 GB at 4-bit
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Qwen3.5 9BAlibaba9.7B≈6.3 GB at 4-bit8 also selling it hosted
- liftDatalab9.7B≈6.3 GB at 4-bit
- MiniCPM-SALAOpenBMB9.5B≈6.2 GB at 4-bit
- AutoGLM-Phone-9B-MultilingualZ.ai9.0B≈5.9 GB at 4-bit3 also selling it hosted
- Llama-3-RefueledRefuel AI8.0B≈5.2 GB at 4-bit
- NuExtract3NuMind4.5B≈3.0 GB at 4-bit
- NuExtractNuMind3.8B≈2.5 GB at 4-bit
- NuExtract-1.5NuMind3.8B≈2.5 GB at 4-bit
- dots.ocrdots studio3.0B≈2.0 GB at 4-bit
- DeepScaleR-1.5B-PreviewAgentica1.8B≈1.2 GB at 4-bit
- MinerU2.5-Pro-2604-1.2BOpenDataLab1.2B≈0.8 GB at 4-bit
- Qwen3.5-0.8BAlibaba0.8B≈0.5 GB at 4-bit1 also selling it hosted
- DolphinByteDance0.4B≈0.3 GB at 4-bit