Structured extraction models that run on 256 GB
21 models with published weights that fit in 256 GB — a Mac Studio with 256. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 274 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A Mac Studio at the top of its configuration, or a serious machine. Almost everything openly published fits, including the large mixtures of experts. At this point the constraint is no longer whether the model loads but whether it generates fast enough to be worth waiting for.
Tools that pull structure out of unstructured input — fields from an invoice, a schema from a page, a table from a report. Some are models, some are services with a model inside. Judge them on what they do when the input is malformed, because that is the whole job; anything can parse a clean file.
- Step 3.7 FlashStepFun201B≈130.9 GB at 4-bit9 also selling it hosted
- Qwen3.5-122B-A10BAlibaba125B≈81.3 GB at 4-bit8 also selling it hosted
- K2-Horizon-MoVA-36B-A4BInstitute of Foundation Models36.0B≈23.4 GB at 4-bit
- Qwen3.5-35B-A3B-BaseAlibaba35.0B≈22.8 GB at 4-bit1 also selling it hosted
- TildeOpen-30bTildeAI30.0B≈19.5 GB at 4-bit
- Qwen3.6 27BAlibaba27.8B≈18.1 GB at 4-bit12 also selling it hosted
- Gemma-3-R1984-12BVIDraft12.2B≈7.9 GB at 4-bit
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Qwen3.5 9BAlibaba9.7B≈6.3 GB at 4-bit8 also selling it hosted
- liftDatalab9.7B≈6.3 GB at 4-bit
- MiniCPM-SALAOpenBMB9.5B≈6.2 GB at 4-bit
- AutoGLM-Phone-9B-MultilingualZ.ai9.0B≈5.9 GB at 4-bit3 also selling it hosted
- Llama-3-RefueledRefuel AI8.0B≈5.2 GB at 4-bit
- NuExtract3NuMind4.5B≈3.0 GB at 4-bit
- NuExtractNuMind3.8B≈2.5 GB at 4-bit
- NuExtract-1.5NuMind3.8B≈2.5 GB at 4-bit
- dots.ocrdots studio3.0B≈2.0 GB at 4-bit
- DeepScaleR-1.5B-PreviewAgentica1.8B≈1.2 GB at 4-bit
- MinerU2.5-Pro-2604-1.2BOpenDataLab1.2B≈0.8 GB at 4-bit
- Qwen3.5-0.8BAlibaba0.8B≈0.5 GB at 4-bit1 also selling it hosted
- DolphinByteDance0.4B≈0.3 GB at 4-bit