Agentic models that run on 256 GB
19 models with published weights that fit in 256 GB — a Mac Studio with 256. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 274 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A Mac Studio at the top of its configuration, or a serious machine. Almost everything openly published fits, including the large mixtures of experts. At this point the constraint is no longer whether the model loads but whether it generates fast enough to be worth waiting for.
- Qwen3 235B A22BAlibaba235B≈152.8 GB at 4-bit10 also selling it hosted
- Qwen3 235B A22B Thinking 2507Alibaba235B≈152.8 GB at 4-bit12 also selling it hosted
- Qwen3-235B-A22B-Instruct-2507Alibaba235B≈152.8 GB at 4-bit13 also selling it hosted
- MiniMax M2.5MiniMax229B≈148.7 GB at 4-bit19 also selling it hosted
- MiniMax-M2MiniMax229B≈148.6 GB at 4-bit12 also selling it hosted
- GPT OSS 120BOpenAI117B≈75.9 GB at 4-bit29 also selling it hosted
- GLM-4.5-AirZ.ai110B≈71.8 GB at 4-bit12 also selling it hosted
- Qwen2.5 VL 72B InstructAlibaba73.4B≈47.7 GB at 4-bit7 also selling it hosted
- Qwen2.5 72B InstructAlibaba72.7B≈47.3 GB at 4-bit12 also selling it hosted
- Qwen2-72B-InstructAlibaba72.0B≈46.8 GB at 4-bit3 also selling it hosted
- Llama 3.3 70B InstructMeta70.6B≈45.9 GB at 4-bit29 also selling it hosted
- Llama-2-70b-chat-hfMeta70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Llama-xLAM-2-70b-fc-rSalesforce70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Meta-Llama-3.1-70B-InstructMeta70.0B≈45.5 GB at 4-bit7 also selling it hosted
- Holo3-35B-A3BH Company35.0B≈22.8 GB at 4-bit1 also selling it hosted
- QwQ-32B-PreviewAlibaba32.0B≈20.8 GB at 4-bit2 also selling it hosted
- xLAM-2-32b-fc-rSalesforce32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen3.5-27BAlibaba27.8B≈18.1 GB at 4-bit8 also selling it hosted
- gpt2-xlOpenAI community1.6B≈1.0 GB at 4-bit