Conversation models that run on 128 GB
613 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 613 in all, a hundred to a page; this is page 1 of 7.
A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- Ling-3.0-flashInclusionAI127B≈82.9 GB at 4-bit7 also selling it hosted
- Meta-Llama-3-120B-Instructmlabonne122B≈79.2 GB at 4-bit
- NVIDIA-Nemotron-3-Super-120B-A12B-FP8NVIDIA120B≈78.0 GB at 4-bit
- galactica-120bfacebook120B≈78.0 GB at 4-bit
- goliath-120balpindale118B≈76.5 GB at 4-bit
- Laguna S 2.1Poolside118B≈76.4 GB at 4-bit4 also selling it hosted
- c4ai-command-a-03-2025Cohere Labs111B≈72.2 GB at 4-bit
- Llama 4 Scout 17BMeta109B≈70.6 GB at 4-bit7 also selling it hosted
- GLM-4.5VZ.ai108B≈70.0 GB at 4-bit9 also selling it hosted
- GLM-4.6VZ.ai108B≈70.0 GB at 4-bit7 also selling it hosted
- sarvam-105bSarvam AI105B≈68.2 GB at 4-bit
- c4ai-command-r-plusCohere Labs104B≈67.5 GB at 4-bit
- Solar-Open-100BUpstage103B≈66.7 GB at 4-bit
- Qwen3 Next 80B A3B InstructAlibaba81.3B≈52.9 GB at 4-bit12 also selling it hosted
- idefics-80b-instructHuggingFaceM480.0B≈52.0 GB at 4-bit
- NVLM-D-72BNVIDIA79.4B≈51.6 GB at 4-bit
- InternVL3-78BOpenGVLab78.4B≈51.0 GB at 4-bit1 also selling it hosted
- calme-3.2-instruct-78bMaziyarPanahi78.0B≈50.7 GB at 4-bit
- Kimi-Dev-72BMoonshot AI72.7B≈47.3 GB at 4-bit
- Magnum v4 72BAnthracite72.7B≈47.3 GB at 4-bit1 also selling it hosted
- Qwen2-72BAlibaba72.7B≈47.3 GB at 4-bit
- Virtuoso LargeArcee AI72.7B≈47.3 GB at 4-bit
- Smaug-72B-v0.1Abacus.AI, Inc.72.3B≈47.0 GB at 4-bit
- Qwen-72BAlibaba72.3B≈47.0 GB at 4-bit
- KAT-Dev-72B-ExpKwaipilot72.0B≈46.8 GB at 4-bit1 also selling it hosted
- Molmo-72B-0924Allen Institute for AI (Ai2)72.0B≈46.8 GB at 4-bit
- QVQ-72B-PreviewAlibaba72.0B≈46.8 GB at 4-bit1 also selling it hosted
- Qwen 2 VL 72B InstructAlibaba72.0B≈46.8 GB at 4-bit4 also selling it hosted
- Qwen-72B-ChatAlibaba72.0B≈46.8 GB at 4-bit
- UI-TARS-72B-DPOByteDance Seed72.0B≈46.8 GB at 4-bit
- dolphin-2.9.2-qwen2-72bdphn72.0B≈46.8 GB at 4-bit1 also selling it hosted
- magnum-v1-72banthracite-org72.0B≈46.8 GB at 4-bit
- Llama3-ChatQA-1.5-70BNVIDIA70.6B≈45.9 GB at 4-bit
- Athene-70BNexusflow70.6B≈45.9 GB at 4-bit
- Hermes 3 70B InstructNous Research70.6B≈45.9 GB at 4-bit4 also selling it hosted
- Hermes 4 70BNous Research70.6B≈45.9 GB at 4-bit4 also selling it hosted
- Higgs-Llama-3-70BBoson AI70.6B≈45.9 GB at 4-bit
- Llama 3.1 Instruct (70B)Meta70.6B≈45.9 GB at 4-bit7 also selling it hosted
- Llama-3.1-Nemotron-70B-Instruct-HFNVIDIA70.6B≈45.9 GB at 4-bit1 also selling it hosted
- Apertus-70B-2509swiss-ai70.0B≈45.5 GB at 4-bit
- Apertus-70B-Instruct-2509swiss-ai70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Apertus-v1.5-70Bswiss-ai70.0B≈45.5 GB at 4-bit1 also selling it hosted
- L3 70B Euryale V2.1Sao10K70.0B≈45.5 GB at 4-bit3 also selling it hosted
- L3.3-70B-Euryale-v2.3Sao10K70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Llama-2-70bMeta Llama70.0B≈45.5 GB at 4-bit
- Llama-2-70b-chatMeta Llama70.0B≈45.5 GB at 4-bit2 also selling it hosted
- Llama-3-Groq-70B-Tool-UseGroq70.0B≈45.5 GB at 4-bit
- Llama-3.1-Nemotron-70B-InstructNVIDIA70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Llama3-OpenBioLLM-70Baaditya70.0B≈45.5 GB at 4-bit
- Platypus2-70B-instructgarage-bAInd70.0B≈45.5 GB at 4-bit
- Reflection-Llama-3.1-70Bmattshumer70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Smaug-Llama-3-70B-InstructAbacus.AI, Inc.70.0B≈45.5 GB at 4-bit
- WizardLM-70B-V1.0WizardLM Team70.0B≈45.5 GB at 4-bit
- Xwin-LM-70B-V0.1Xwin-LM70.0B≈45.5 GB at 4-bit
- dolphin-2.9.1-llama-3-70bcognitivecomputations70.0B≈45.5 GB at 4-bit1 also selling it hosted
- llama-3-70B-Instruct-abliteratedfailspy70.0B≈45.5 GB at 4-bit
- lzlv_70b_fp16_hflizpreciatior70.0B≈45.5 GB at 4-bit1 also selling it hosted
- med42-70bM42 Health70.0B≈45.5 GB at 4-bit
- meditron-70bEPFL LLM Team70.0B≈45.5 GB at 4-bit
- tulu-2-dpo-70bAllen Institute for AI (Ai2)70.0B≈45.5 GB at 4-bit
- LongCat-Flash-LiteMeituan LongCat69.1B≈44.9 GB at 4-bit
- Llama-2-70b-hfMeta Llama69.0B≈44.8 GB at 4-bit
- NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4NVIDIA67.2B≈43.7 GB at 4-bit
- opt-66bfacebook66.0B≈42.9 GB at 4-bit
- Jamba-v0.1AI21 Labs51.6B≈33.5 GB at 4-bit
- Llama-3_1-Nemotron-51B-InstructNVIDIA51.5B≈33.5 GB at 4-bit
- Kimi-Linear-48B-A3B-InstructMoonshot AI49.1B≈31.9 GB at 4-bit
- Llama-3_3-Nemotron-Super-49B-v1NVIDIA49.0B≈31.9 GB at 4-bit
- dolphin-2.5-mixtral-8x7bDolphin46.7B≈30.4 GB at 4-bit
- GRIN-MoEMicrosoft41.9B≈27.2 GB at 4-bit
- Phi-3.5-MoE-instructMicrosoft41.9B≈27.2 GB at 4-bit1 also selling it hosted
- falcon-40bTechnology Innovation Institute41.8B≈27.2 GB at 4-bit
- falcon-40b-instructTechnology Innovation Institute40.0B≈26.0 GB at 4-bit
- CLIP-ViT-bigG-14-laion2B-39B-b160kLAION eV39.0B≈25.4 GB at 4-bit
- Skyfall 36B V2TheDrummer36.0B≈23.4 GB at 4-bit1 also selling it hosted
- BigBang-v1The Endless Frontier36.0B≈23.4 GB at 4-bit
- Huihui-Qwen3.5-35B-A3B-abliteratedhuihui-ai36.0B≈23.4 GB at 4-bit
- Ornith-1.5-35B-A3BOrnith36.0B≈23.4 GB at 4-bit
- Agents-A1Intern Science35.1B≈22.8 GB at 4-bit
- Nex-N2-MiniNEX AGI35.1B≈22.8 GB at 4-bit
- Nex-N2.5-miniNEX AGI35.1B≈22.8 GB at 4-bit1 also selling it hosted
- Thomson-1.0-SmallThomson Reuters35.1B≈22.8 GB at 4-bit
- XYZ-Aquila-miniXYZAILab35.1B≈22.8 GB at 4-bit
- Ornith-1.0-35Bdeepreinforce-ai35.0B≈22.8 GB at 4-bit1 also selling it hosted
- Qwen3.6-35B-A3B-DFlashZ Lab35.0B≈22.8 GB at 4-bit
- aya-23-35BCohere Labs35.0B≈22.8 GB at 4-bit
- c4ai-command-r-v01Cohere Labs35.0B≈22.7 GB at 4-bit
- llava-v1.6-34bliuhaotian34.8B≈22.6 GB at 4-bit
- Qwen-AgentWorld-35B-A3BAlibaba34.7B≈22.5 GB at 4-bit
- Yi-34B01-ai34.4B≈22.4 GB at 4-bit1 also selling it hosted
- Ovis2-34BATH-MaaS34.0B≈22.1 GB at 4-bit
- Yi-VL-34B01-ai34.0B≈22.1 GB at 4-bit
- deepsex-34bTriadParty34.0B≈22.1 GB at 4-bit
- Laguna XS 2.1Poolside33.4B≈21.7 GB at 4-bit3 also selling it hosted
- Qwen3 VL 32B InstructAlibaba33.4B≈21.7 GB at 4-bit5 also selling it hosted
- HyperCLOVAX-SEED-Think-32BHyperCLOVA X33.3B≈21.7 GB at 4-bit
- aya-vision-32bCohere Labs33.1B≈21.5 GB at 4-bit
- EXAONE-4.5-33BLGAI-EXAONE33.0B≈21.4 GB at 4-bit
- DeepSeek-R1-Distill-Qwen-32B-JapaneseCyberAgent32.8B≈21.3 GB at 4-bit
- K2-ThinkInstitute of Foundation Models32.8B≈21.3 GB at 4-bit