Evaluation models that run on 32 GB
16 models with published weights that fit in 32 GB — a well-specified laptop. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 33 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
Enough for a thirty-billion-parameter model at four-bit with a long context, or a smaller one at higher precision if quality matters more than size. A practical ceiling for a laptop that also has to be a laptop.
Tools for judging what a model produced: scoring runs, tracing a chain of calls, catching a regression before a user does. Priced per trace or per evaluation. Most use a model as the judge, which is worth knowing, because it means your evaluation has a bill and an opinion of its own.
- Olmo 3 32B ThinkAllen Institute for AI (Ai2)32.2B≈21.0 GB at 4-bit
- Qwen3 30B A3B Thinking 2507Alibaba30.5B≈19.8 GB at 4-bit4 also selling it hosted
- DeepSeek-Coder-V2-Lite-InstructDeepSeek15.7B≈10.2 GB at 4-bit1 also selling it hosted
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Qwythos-9B-Claude-Mythos-5-1Mempero-ai9.4B≈6.1 GB at 4-bit
- Nemotron Nano 9B V2NVIDIA9.0B≈5.9 GB at 4-bit5 also selling it hosted
- Aion-RP 1.0 (8B)AionLabs8.0B≈5.2 GB at 4-bit2 also selling it hosted
- DeepSeek R1 0528 Qwen3 8BDeepSeek8.0B≈5.2 GB at 4-bit1 also selling it hosted
- DeepSeek-R1-Distill-Llama-8BDeepSeek8.0B≈5.2 GB at 4-bit3 also selling it hosted
- Qwen2.5-Omni-3BAlibaba5.5B≈3.6 GB at 4-bit
- fable-tracesAliesTaha4.0B≈2.6 GB at 4-bit
- LocoOperator-4BLocoreMind4.0B≈2.6 GB at 4-bit
- Isaac 0.2 2B PreviewPerceptron2.0B≈1.3 GB at 4-bit1 also selling it hosted
- Isaac 0.2 1BPerceptron1.0B≈0.7 GB at 4-bit1 also selling it hosted
- llama-nemotron-rerank-vl-1b-v2NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- PaddleOCR-VL-0.9Bpaddlepaddle0.9B≈0.6 GB at 4-bit1 also selling it hosted