Reasoning models that run on 16 GB
33 models with published weights that fit in 16 GB — the common laptop. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 16 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
The commonest laptop configuration, and the point where a genuinely useful model runs locally: fourteen billion parameters at four-bit fits with room to spare for the context. Expect a capable general assistant that will not match the frontier, and remember that anything else the machine is doing competes for the same memory.
Models that plan before they answer, spending extra tokens on working the problem through. They win on mathematics, hard code and anything with several steps, and they are measured on boards like GPQA, AIME and ARC-AGI. The catch is what the thinking costs: reasoning is billed as output tokens, so the same model can cost several times more per answer at a high effort than a low one, and for a question that needed one turn you have paid for a monologue.
- Phi-4-reasoning-vision-15BMicrosoft15.0B≈9.8 GB at 4-bit
- Qwen3 14BAlibaba14.8B≈9.6 GB at 4-bit8 also selling it hosted
- Phi-4Microsoft14.7B≈9.5 GB at 4-bit6 also selling it hosted
- DeepSeek R1 Distill QWEN 14BDeepSeek14.0B≈9.1 GB at 4-bit5 also selling it hosted
- Phi-3-medium-128k-instructMicrosoft14.0B≈9.1 GB at 4-bit1 also selling it hosted
- Gemma 3 12BGoogle12.2B≈7.9 GB at 4-bit9 also selling it hosted
- Mellum2-12B-A2.5B-ThinkingJetBrains12.1B≈7.9 GB at 4-bit
- Qwen3.5 9BAlibaba9.7B≈6.3 GB at 4-bit8 also selling it hosted
- Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSOREDDavidAU9.4B≈6.1 GB at 4-bit
- gemma-2-9b-itGoogle9.0B≈5.9 GB at 4-bit3 also selling it hosted
- Qwen3 VL 8B ThinkingAlibaba8.8B≈5.7 GB at 4-bit3 also selling it hosted
- NuMarkdown-8B-ThinkingNuMind8.3B≈5.4 GB at 4-bit
- Qwen3 8BAlibaba8.2B≈5.3 GB at 4-bit8 also selling it hosted
- Llama 3.1 8B InstructMeta8.0B≈5.2 GB at 4-bit20 also selling it hosted
- DeepSeek R1 0528 Qwen3 8BDeepSeek8.0B≈5.2 GB at 4-bit1 also selling it hosted
- Seed-Coder-8B-ReasoningByteDance Seed8.0B≈5.2 GB at 4-bit
- Qwen2.5 7B InstructAlibaba7.6B≈5.0 GB at 4-bit8 also selling it hosted
- Mistral-7B-Instruct-v0.3Mistral AI7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Gemma 3 4BGoogle4.3B≈2.8 GB at 4-bit6 also selling it hosted
- DASD-4B-ThinkingAlibaba Cloud Apsara Lab4.0B≈2.6 GB at 4-bit
- Qwen3 4BAlibaba4.0B≈2.6 GB at 4-bit3 also selling it hosted
- Qwen3-4B-Instruct-2507Alibaba4.0B≈2.6 GB at 4-bit2 also selling it hosted
- Qwen3-4B-Thinking-2507Alibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3-VL-4B-ThinkingAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3.5-4BAlibaba4.0B≈2.6 GB at 4-bit2 also selling it hosted
- Granite 4.0 MicroIBM watsonx.ai3.2B≈2.1 GB at 4-bit2 also selling it hosted
- Ouro-2.6B-ThinkingByteDance2.6B≈1.7 GB at 4-bit
- Qwen3-1.7BAlibaba2.0B≈1.3 GB at 4-bit1 also selling it hosted
- DeepSeek-R1-Distill-Qwen-1.5BDeepSeek1.8B≈1.2 GB at 4-bit3 also selling it hosted
- Nemotron-Research-Reasoning-Qwen-1.5BNVIDIA1.8B≈1.2 GB at 4-bit
- Llama 3.2 1B InstructMeta1.2B≈0.8 GB at 4-bit11 also selling it hosted
- LFM2.5-1.2B-ThinkingLiquid AI1.2B≈0.8 GB at 4-bit
- MiniCPM5-1B-Claude-Opus-Fable5-ThinkingGnLOLot1.0B≈0.7 GB at 4-bit