Agentic models that run on 64 GB
5 models with published weights that fit in 64 GB — a workstation. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 67 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A workstation, or a well-specified Mac. Seventy-billion-parameter models fit at four-bit, which is the class where local output stops being obviously worse than the hosted models people pay for. Loading takes real time and the machine will be warm.
- Holo3-35B-A3BH Company35.0B≈22.8 GB at 4-bit1 also selling it hosted
- QwQ-32B-PreviewAlibaba32.0B≈20.8 GB at 4-bit2 also selling it hosted
- xLAM-2-32b-fc-rSalesforce32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen3.5-27BAlibaba27.8B≈18.1 GB at 4-bit8 also selling it hosted
- gpt2-xlOpenAI community1.6B≈1.0 GB at 4-bit