Models that run on 128 GB
1133 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 1,133 in all, a hundred to a page; this is page 2 of 12.
A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.
- Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingDavidAU39.5B≈25.7 GB at 4-bit
- CLIP-ViT-bigG-14-laion2B-39B-b160kLAION eV39.0B≈25.4 GB at 4-bit
- Seed-OSS-36B-InstructByteDance Seed36.2B≈23.5 GB at 4-bit1 also selling it hosted
- K2-Horizon-MoVA-36B-A4BInstitute of Foundation Models36.0B≈23.4 GB at 4-bit
- Skyfall 36B V2TheDrummer36.0B≈23.4 GB at 4-bit1 also selling it hosted
- BigBang-v1The Endless Frontier36.0B≈23.4 GB at 4-bit
- Huihui-Qwen3.5-35B-A3B-abliteratedhuihui-ai36.0B≈23.4 GB at 4-bit
- Ornith-1.5-35B-A3BOrnith36.0B≈23.4 GB at 4-bit
- Qwable-v1lordx6436.0B≈23.4 GB at 4-bit
- Qwen3.5-35B-A3BAlibaba36.0B≈23.4 GB at 4-bit9 also selling it hosted
- Qwen3.6 35B A3BAlibaba36.0B≈23.4 GB at 4-bit10 also selling it hosted
- Agents-A1Intern Science35.1B≈22.8 GB at 4-bit
- Nex-N2-MiniNEX AGI35.1B≈22.8 GB at 4-bit
- Nex-N2.5-miniNEX AGI35.1B≈22.8 GB at 4-bit1 also selling it hosted
- Thomson-1.0-SmallThomson Reuters35.1B≈22.8 GB at 4-bit
- XYZ-Aquila-miniXYZAILab35.1B≈22.8 GB at 4-bit
- Holo3-35B-A3BH Company35.0B≈22.8 GB at 4-bit1 also selling it hosted
- Ornith-1.0-35Bdeepreinforce-ai35.0B≈22.8 GB at 4-bit1 also selling it hosted
- Qwen3.5-35B-A3B-BaseAlibaba35.0B≈22.8 GB at 4-bit1 also selling it hosted
- Qwen3.6-35B-A3B-DFlashZ Lab35.0B≈22.8 GB at 4-bit
- aya-23-35BCohere Labs35.0B≈22.8 GB at 4-bit
- c4ai-command-r-v01Cohere Labs35.0B≈22.7 GB at 4-bit
- llava-v1.6-34bliuhaotian34.8B≈22.6 GB at 4-bit
- KAT-Coder-V2.5-DevKwaipilot34.7B≈22.5 GB at 4-bit
- Qwen-AgentWorld-35B-A3BAlibaba34.7B≈22.5 GB at 4-bit
- Yi-34B01-ai34.4B≈22.4 GB at 4-bit1 also selling it hosted
- Yi-34B-Chat01-ai34.4B≈22.4 GB at 4-bit1 also selling it hosted
- CodeBooga-34B-v0.1oobabooga34.0B≈22.1 GB at 4-bit
- CodeLlama-34b-Instruct-hfCode Llama34.0B≈22.1 GB at 4-bit
- CodeLlama-34b-hfCode Llama34.0B≈22.1 GB at 4-bit
- Ovis2-34BATH-MaaS34.0B≈22.1 GB at 4-bit
- Phind-CodeLlama-34B-Python-v1Phind34.0B≈22.1 GB at 4-bit1 also selling it hosted
- Phind-CodeLlama-34B-v1Phind34.0B≈22.1 GB at 4-bit1 also selling it hosted
- Phind-CodeLlama-34B-v2Phind34.0B≈22.1 GB at 4-bit2 also selling it hosted
- WizardCoder-Python-34B-V1.0WizardLM Team34.0B≈22.1 GB at 4-bit
- Yi-1.5-34B-Chat01-ai34.0B≈22.1 GB at 4-bit
- Yi-VL-34B01-ai34.0B≈22.1 GB at 4-bit
- deepsex-34bTriadParty34.0B≈22.1 GB at 4-bit
- sqlcoder-34b-alphadefog34.0B≈22.1 GB at 4-bit
- Laguna XS 2.1Poolside33.4B≈21.7 GB at 4-bit3 also selling it hosted
- Qwen3 VL 32B InstructAlibaba33.4B≈21.7 GB at 4-bit5 also selling it hosted
- deepseek-coder-33b-instructDeepSeek33.3B≈21.7 GB at 4-bit2 also selling it hosted
- HyperCLOVAX-SEED-Think-32BHyperCLOVA X33.3B≈21.7 GB at 4-bit
- aya-vision-32bCohere Labs33.1B≈21.5 GB at 4-bit
- Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16NVIDIA33.0B≈21.5 GB at 4-bit
- EXAONE-4.5-33BLGAI-EXAONE33.0B≈21.4 GB at 4-bit
- OpenCodeInterpreter-DS-33BMultimodal Art Projection33.0B≈21.4 GB at 4-bit
- Qwen3 32BAlibaba32.8B≈21.3 GB at 4-bit17 also selling it hosted
- AM-Thinking-v1am team32.8B≈21.3 GB at 4-bit
- DeepSeek-R1-Distill-Qwen-32B-JapaneseCyberAgent32.8B≈21.3 GB at 4-bit
- K2-ThinkInstitute of Foundation Models32.8B≈21.3 GB at 4-bit
- Qwen2.5 Coder 32B InstructAlibaba32.8B≈21.3 GB at 4-bit11 also selling it hosted
- TinyR1-32B-Previewqihoo36032.8B≈21.3 GB at 4-bit
- openhands-lm-32b-v0.1OpenHands32.8B≈21.3 GB at 4-bit
- DeepSWE-PreviewAgentica32.8B≈21.3 GB at 4-bit
- KAT-DevKwaipilot32.8B≈21.3 GB at 4-bit
- GLM-Z1-32B-0414zai-org32.6B≈21.2 GB at 4-bit
- Olmo 3 32B ThinkAllen Institute for AI (Ai2)32.2B≈21.0 GB at 4-bit
- FLUX.2 [dev]Black Forest Labs32.2B≈20.9 GB at 4-bit1 also selling it hosted
- sarvam-30bSarvam AI32.2B≈20.9 GB at 4-bit
- Baichuan M2 32Bbaichuan32.0B≈20.8 GB at 4-bit2 also selling it hosted
- DeepSeek R1 Distill QWEN 32BDeepSeek32.0B≈20.8 GB at 4-bit6 also selling it hosted
- EXAONE-Deep-32BLGAI-EXAONE32.0B≈20.8 GB at 4-bit
- GLM-4-32B-0414Z.ai32.0B≈20.8 GB at 4-bit3 also selling it hosted
- Olmo-2-0325-32B-InstructAllen Institute for AI (Ai2)32.0B≈20.8 GB at 4-bit
- Olmo-3.1-32B-InstructAllen Institute for AI (Ai2)32.0B≈20.8 GB at 4-bit1 also selling it hosted
- OlympicCoder-32Bopen-r132.0B≈20.8 GB at 4-bit
- OpenReasoning-Nemotron-32BNVIDIA32.0B≈20.8 GB at 4-bit1 also selling it hosted
- OpenThinker-32Bopen-thoughts32.0B≈20.8 GB at 4-bit
- QwQ-32BAlibaba32.0B≈20.8 GB at 4-bit7 also selling it hosted
- QwQ-32B-PreviewAlibaba32.0B≈20.8 GB at 4-bit2 also selling it hosted
- QwQ-R1984-32BVIDraft32.0B≈20.8 GB at 4-bit
- Qwen-SEA-LION-v4-32B-ITaisingapore32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen2.5-32BAlibaba32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen2.5-VL-32B-InstructAlibaba32.0B≈20.8 GB at 4-bit3 also selling it hosted
- QwenLong-L1-32BTongyi-Zhiwen32.0B≈20.8 GB at 4-bit
- Sky-T1-32B-PreviewNovaSky-AI32.0B≈20.8 GB at 4-bit1 also selling it hosted
- UIGEN-X-32B-0727Tesslate32.0B≈20.8 GB at 4-bit
- s1-32Bsimplescaling32.0B≈20.8 GB at 4-bit
- xLAM-2-32b-fc-rSalesforce32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen3-Omni-30B-A3B-CaptionerAlibaba31.7B≈20.6 GB at 4-bit
- NVIDIA Nemotron 3.5 Lightning 30B A3BNVIDIA31.6B≈20.5 GB at 4-bit2 also selling it hosted
- Nemotron 3 Nano 30B A3BNVIDIA31.6B≈20.5 GB at 4-bit10 also selling it hosted
- Nemotron-Cascade-2-30B-A3BNVIDIA31.6B≈20.5 GB at 4-bit
- phonellm-alpha-1Pipecat31.6B≈20.5 GB at 4-bit
- Gemma 4 31BGoogle31.3B≈20.3 GB at 4-bit15 also selling it hosted
- Qwen3 VL 30B A3B InstructAlibaba31.1B≈20.2 GB at 4-bit8 also selling it hosted
- Qwen3 VL 30B A3B ThinkingAlibaba31.1B≈20.2 GB at 4-bit7 also selling it hosted
- Gemma-4-31B-it-assistantGoogle31.0B≈20.2 GB at 4-bit
- Gemma-4-31B-it-pearlpearl-ai31.0B≈20.2 GB at 4-bit
- Qwen3-Coder-30B-A3B-Instruct-FP8Alibaba30.5B≈19.8 GB at 4-bit
- MiroThinker-v1.5-30BMiroMind AI30.5B≈19.8 GB at 4-bit
- Qwen3 30B A3BAlibaba30.5B≈19.8 GB at 4-bit11 also selling it hosted
- Qwen3 30B A3B Instruct 2507Alibaba30.5B≈19.8 GB at 4-bit8 also selling it hosted
- Qwen3 30B A3B Thinking 2507Alibaba30.5B≈19.8 GB at 4-bit4 also selling it hosted
- Qwen3 Coder 30B A3B InstructAlibaba30.5B≈19.8 GB at 4-bit11 also selling it hosted
- Tongyi-DeepResearch-30B-A3BAlibaba-NLP30.5B≈19.8 GB at 4-bit
- Hy-MT2-30B-A3BTencent Hunyuan30.0B≈19.5 GB at 4-bit1 also selling it hosted
- Nemotron-Labs-Audex-30B-A3BNVIDIA30.0B≈19.5 GB at 4-bit
- Ovis2.6-30B-A3BATH-MaaS30.0B≈19.5 GB at 4-bit