Agentic models that run on 128 GB
14 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.
- GPT OSS 120BOpenAI117B≈75.9 GB at 4-bit29 also selling it hosted
- GLM-4.5-AirZ.ai110B≈71.8 GB at 4-bit12 also selling it hosted
- Qwen2.5 VL 72B InstructAlibaba73.4B≈47.7 GB at 4-bit7 also selling it hosted
- Qwen2.5 72B InstructAlibaba72.7B≈47.3 GB at 4-bit12 also selling it hosted
- Qwen2-72B-InstructAlibaba72.0B≈46.8 GB at 4-bit3 also selling it hosted
- Llama 3.3 70B InstructMeta70.6B≈45.9 GB at 4-bit29 also selling it hosted
- Llama-2-70b-chat-hfMeta70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Llama-xLAM-2-70b-fc-rSalesforce70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Meta-Llama-3.1-70B-InstructMeta70.0B≈45.5 GB at 4-bit7 also selling it hosted
- Holo3-35B-A3BH Company35.0B≈22.8 GB at 4-bit1 also selling it hosted
- QwQ-32B-PreviewAlibaba32.0B≈20.8 GB at 4-bit2 also selling it hosted
- xLAM-2-32b-fc-rSalesforce32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen3.5-27BAlibaba27.8B≈18.1 GB at 4-bit8 also selling it hosted
- gpt2-xlOpenAI community1.6B≈1.0 GB at 4-bit