Conversation models that run on 64 GB
551 models with published weights that fit in 64 GB — a workstation. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 67 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 551 in all, a hundred to a page; this is page 1 of 6.
A workstation, or a well-specified Mac. Seventy-billion-parameter models fit at four-bit, which is the class where local output stops being obviously worse than the hosted models people pay for. Loading takes real time and the machine will be warm.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4NVIDIA67.2B≈43.7 GB at 4-bit
- opt-66bfacebook66.0B≈42.9 GB at 4-bit
- Jamba-v0.1AI21 Labs51.6B≈33.5 GB at 4-bit
- Llama-3_1-Nemotron-51B-InstructNVIDIA51.5B≈33.5 GB at 4-bit
- Kimi-Linear-48B-A3B-InstructMoonshot AI49.1B≈31.9 GB at 4-bit
- Llama-3_3-Nemotron-Super-49B-v1NVIDIA49.0B≈31.9 GB at 4-bit
- dolphin-2.5-mixtral-8x7bDolphin46.7B≈30.4 GB at 4-bit
- GRIN-MoEMicrosoft41.9B≈27.2 GB at 4-bit
- Phi-3.5-MoE-instructMicrosoft41.9B≈27.2 GB at 4-bit1 also selling it hosted
- falcon-40bTechnology Innovation Institute41.8B≈27.2 GB at 4-bit
- falcon-40b-instructTechnology Innovation Institute40.0B≈26.0 GB at 4-bit
- CLIP-ViT-bigG-14-laion2B-39B-b160kLAION eV39.0B≈25.4 GB at 4-bit
- Skyfall 36B V2TheDrummer36.0B≈23.4 GB at 4-bit1 also selling it hosted
- BigBang-v1The Endless Frontier36.0B≈23.4 GB at 4-bit
- Huihui-Qwen3.5-35B-A3B-abliteratedhuihui-ai36.0B≈23.4 GB at 4-bit
- Ornith-1.5-35B-A3BOrnith36.0B≈23.4 GB at 4-bit
- Agents-A1Intern Science35.1B≈22.8 GB at 4-bit
- Nex-N2-MiniNEX AGI35.1B≈22.8 GB at 4-bit
- Nex-N2.5-miniNEX AGI35.1B≈22.8 GB at 4-bit1 also selling it hosted
- Thomson-1.0-SmallThomson Reuters35.1B≈22.8 GB at 4-bit
- XYZ-Aquila-miniXYZAILab35.1B≈22.8 GB at 4-bit
- Ornith-1.0-35Bdeepreinforce-ai35.0B≈22.8 GB at 4-bit1 also selling it hosted
- Qwen3.6-35B-A3B-DFlashZ Lab35.0B≈22.8 GB at 4-bit
- aya-23-35BCohere Labs35.0B≈22.8 GB at 4-bit
- c4ai-command-r-v01Cohere Labs35.0B≈22.7 GB at 4-bit
- llava-v1.6-34bliuhaotian34.8B≈22.6 GB at 4-bit
- Qwen-AgentWorld-35B-A3BAlibaba34.7B≈22.5 GB at 4-bit
- Yi-34B01-ai34.4B≈22.4 GB at 4-bit1 also selling it hosted
- Ovis2-34BATH-MaaS34.0B≈22.1 GB at 4-bit
- Yi-VL-34B01-ai34.0B≈22.1 GB at 4-bit
- deepsex-34bTriadParty34.0B≈22.1 GB at 4-bit
- Laguna XS 2.1Poolside33.4B≈21.7 GB at 4-bit3 also selling it hosted
- Qwen3 VL 32B InstructAlibaba33.4B≈21.7 GB at 4-bit5 also selling it hosted
- HyperCLOVAX-SEED-Think-32BHyperCLOVA X33.3B≈21.7 GB at 4-bit
- aya-vision-32bCohere Labs33.1B≈21.5 GB at 4-bit
- EXAONE-4.5-33BLGAI-EXAONE33.0B≈21.4 GB at 4-bit
- DeepSeek-R1-Distill-Qwen-32B-JapaneseCyberAgent32.8B≈21.3 GB at 4-bit
- K2-ThinkInstitute of Foundation Models32.8B≈21.3 GB at 4-bit
- TinyR1-32B-Previewqihoo36032.8B≈21.3 GB at 4-bit
- DeepSWE-PreviewAgentica32.8B≈21.3 GB at 4-bit
- KAT-DevKwaipilot32.8B≈21.3 GB at 4-bit
- GLM-Z1-32B-0414zai-org32.6B≈21.2 GB at 4-bit
- sarvam-30bSarvam AI32.2B≈20.9 GB at 4-bit
- Baichuan M2 32Bbaichuan32.0B≈20.8 GB at 4-bit2 also selling it hosted
- EXAONE-Deep-32BLGAI-EXAONE32.0B≈20.8 GB at 4-bit
- GLM-4-32B-0414Z.ai32.0B≈20.8 GB at 4-bit3 also selling it hosted
- Olmo-2-0325-32B-InstructAllen Institute for AI (Ai2)32.0B≈20.8 GB at 4-bit
- Olmo-3.1-32B-InstructAllen Institute for AI (Ai2)32.0B≈20.8 GB at 4-bit1 also selling it hosted
- OpenThinker-32Bopen-thoughts32.0B≈20.8 GB at 4-bit
- QwQ-R1984-32BVIDraft32.0B≈20.8 GB at 4-bit
- Qwen-SEA-LION-v4-32B-ITaisingapore32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen2.5-32BAlibaba32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen2.5-VL-32B-InstructAlibaba32.0B≈20.8 GB at 4-bit3 also selling it hosted
- QwenLong-L1-32BTongyi-Zhiwen32.0B≈20.8 GB at 4-bit
- Sky-T1-32B-PreviewNovaSky-AI32.0B≈20.8 GB at 4-bit1 also selling it hosted
- UIGEN-X-32B-0727Tesslate32.0B≈20.8 GB at 4-bit
- s1-32Bsimplescaling32.0B≈20.8 GB at 4-bit
- NVIDIA Nemotron 3.5 Lightning 30B A3BNVIDIA31.6B≈20.5 GB at 4-bit2 also selling it hosted
- Nemotron 3 Nano 30B A3BNVIDIA31.6B≈20.5 GB at 4-bit10 also selling it hosted
- Nemotron-Cascade-2-30B-A3BNVIDIA31.6B≈20.5 GB at 4-bit
- phonellm-alpha-1Pipecat31.6B≈20.5 GB at 4-bit
- Qwen3 VL 30B A3B InstructAlibaba31.1B≈20.2 GB at 4-bit8 also selling it hosted
- Gemma-4-31B-it-pearlpearl-ai31.0B≈20.2 GB at 4-bit
- MiroThinker-v1.5-30BMiroMind AI30.5B≈19.8 GB at 4-bit
- Tongyi-DeepResearch-30B-A3BAlibaba-NLP30.5B≈19.8 GB at 4-bit
- Nemotron-Labs-Audex-30B-A3BNVIDIA30.0B≈19.5 GB at 4-bit
- Ovis2.6-30B-A3BATH-MaaS30.0B≈19.5 GB at 4-bit
- QwenLong-L1.5-30B-A3BTongyi-Zhiwen30.0B≈19.5 GB at 4-bit
- Wizard-Vicuna-30B-UncensoredQuixi AI30.0B≈19.5 GB at 4-bit
- WizardLM-30B-UncensoredQuixi AI30.0B≈19.5 GB at 4-bit
- granite-4.1-30bIBM Granite30.0B≈19.5 GB at 4-bit
- Muse Glimmer 30BMeta29.8B≈19.4 GB at 4-bit8 also selling it hosted
- medgemma-27b-itGoogle28.8B≈18.7 GB at 4-bit
- ERNIE 4.5 VL 28B A3BBaidu28.0B≈18.2 GB at 4-bit2 also selling it hosted
- Huihui-Qwen3.8-27B-abliteratedhuihui-ai27.8B≈18.1 GB at 4-bit
- Qwen3.8 27BAlibaba27.8B≈18.1 GB at 4-bit12 also selling it hosted
- Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16AEON-727.8B≈18.1 GB at 4-bit
- Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAUDavidAU27.8B≈18.1 GB at 4-bit
- Qwen3.8-27B-UncensoredOrcaRouter27.8B≈18.1 GB at 4-bit
- Swift-Qwen3.8-27bukisai27.8B≈18.1 GB at 4-bit
- deepseek-vl2DeepSeek27.5B≈17.9 GB at 4-bit
- Gemma-3-R1984-27BVIDraft27.4B≈17.8 GB at 4-bit
- gemma-3-27b-it-abliteratedmlabonne27.4B≈17.8 GB at 4-bit
- datagemma-rag-27b-itGoogle27.2B≈17.7 GB at 4-bit
- medgemma-27b-text-itGoogle27.0B≈17.6 GB at 4-bit
- C2S-Scale-Gemma-2-27Bvandijklab27.0B≈17.6 GB at 4-bit
- Fara1.5-27BMicrosoft27.0B≈17.6 GB at 4-bit
- Gemma-SEA-LION-v4-27B-ITaisingapore27.0B≈17.6 GB at 4-bit2 also selling it hosted
- Huihui-Qwen3.5-27B-abliteratedhuihui-ai27.0B≈17.6 GB at 4-bit
- Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16AEON-727.0B≈17.6 GB at 4-bit
- Qwen3.8-27B-DFlash2Z Lab27.0B≈17.6 GB at 4-bit
- Qwythos-27B-v1empero-ai27.0B≈17.6 GB at 4-bit
- Gemma 4 26B A4B ITGoogle26.0B≈16.9 GB at 4-bit
- diffusiongemma-26B-A4B-itGoogle25.8B≈16.8 GB at 4-bit
- Gemma 4 26B A4BGoogle25.8B≈16.8 GB at 4-bit7 also selling it hosted
- AriaRhymes.AI25.3B≈16.4 GB at 4-bit
- Dolphin Mistral 24B Venice EditionCognitive Computations24.0B≈15.6 GB at 4-bit1 also selling it hosted
- Mistral Small 3.1 24B Instruct 2503Mistral AI24.0B≈15.6 GB at 4-bit3 also selling it hosted
- Mistral-Small-3.2-24B-Instruct-2506Mistral AI24.0B≈15.6 GB at 4-bit1 also selling it hosted
- Mistral Small 3Mistral AI23.6B≈15.3 GB at 4-bit4 also selling it hosted