Models that run on 96 GB
1115 models with published weights that fit in 96 GB — a MacBook Pro with 96. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 102 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 1,115 in all, a hundred to a page; this is page 2 of 12.
An Apple configuration for people who intend to run models rather than occasionally try one. Comfortably holds the seventy-billion class with a very long context, or a mixture-of-experts model whose active parameters are few but whose weights are all resident.
- Qwen3.5-35B-A3B-BaseAlibaba35.0B≈22.8 GB at 4-bit1 also selling it hosted
- Qwen3.6-35B-A3B-DFlashZ Lab35.0B≈22.8 GB at 4-bit
- aya-23-35BCohere Labs35.0B≈22.8 GB at 4-bit
- c4ai-command-r-v01Cohere Labs35.0B≈22.7 GB at 4-bit
- llava-v1.6-34bliuhaotian34.8B≈22.6 GB at 4-bit
- KAT-Coder-V2.5-DevKwaipilot34.7B≈22.5 GB at 4-bit
- Qwen-AgentWorld-35B-A3BAlibaba34.7B≈22.5 GB at 4-bit
- Yi-34B01-ai34.4B≈22.4 GB at 4-bit1 also selling it hosted
- Yi-34B-Chat01-ai34.4B≈22.4 GB at 4-bit1 also selling it hosted
- CodeBooga-34B-v0.1oobabooga34.0B≈22.1 GB at 4-bit
- CodeLlama-34b-Instruct-hfCode Llama34.0B≈22.1 GB at 4-bit
- CodeLlama-34b-hfCode Llama34.0B≈22.1 GB at 4-bit
- Ovis2-34BATH-MaaS34.0B≈22.1 GB at 4-bit
- Phind-CodeLlama-34B-Python-v1Phind34.0B≈22.1 GB at 4-bit1 also selling it hosted
- Phind-CodeLlama-34B-v1Phind34.0B≈22.1 GB at 4-bit1 also selling it hosted
- Phind-CodeLlama-34B-v2Phind34.0B≈22.1 GB at 4-bit2 also selling it hosted
- WizardCoder-Python-34B-V1.0WizardLM Team34.0B≈22.1 GB at 4-bit
- Yi-1.5-34B-Chat01-ai34.0B≈22.1 GB at 4-bit
- Yi-VL-34B01-ai34.0B≈22.1 GB at 4-bit
- deepsex-34bTriadParty34.0B≈22.1 GB at 4-bit
- sqlcoder-34b-alphadefog34.0B≈22.1 GB at 4-bit
- Laguna XS 2.1Poolside33.4B≈21.7 GB at 4-bit3 also selling it hosted
- Qwen3 VL 32B InstructAlibaba33.4B≈21.7 GB at 4-bit5 also selling it hosted
- deepseek-coder-33b-instructDeepSeek33.3B≈21.7 GB at 4-bit2 also selling it hosted
- HyperCLOVAX-SEED-Think-32BHyperCLOVA X33.3B≈21.7 GB at 4-bit
- aya-vision-32bCohere Labs33.1B≈21.5 GB at 4-bit
- Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16NVIDIA33.0B≈21.5 GB at 4-bit
- EXAONE-4.5-33BLGAI-EXAONE33.0B≈21.4 GB at 4-bit
- OpenCodeInterpreter-DS-33BMultimodal Art Projection33.0B≈21.4 GB at 4-bit
- Qwen3 32BAlibaba32.8B≈21.3 GB at 4-bit17 also selling it hosted
- AM-Thinking-v1am team32.8B≈21.3 GB at 4-bit
- DeepSeek-R1-Distill-Qwen-32B-JapaneseCyberAgent32.8B≈21.3 GB at 4-bit
- K2-ThinkInstitute of Foundation Models32.8B≈21.3 GB at 4-bit
- Qwen2.5 Coder 32B InstructAlibaba32.8B≈21.3 GB at 4-bit11 also selling it hosted
- TinyR1-32B-Previewqihoo36032.8B≈21.3 GB at 4-bit
- openhands-lm-32b-v0.1OpenHands32.8B≈21.3 GB at 4-bit
- DeepSWE-PreviewAgentica32.8B≈21.3 GB at 4-bit
- KAT-DevKwaipilot32.8B≈21.3 GB at 4-bit
- GLM-Z1-32B-0414zai-org32.6B≈21.2 GB at 4-bit
- Olmo 3 32B ThinkAllen Institute for AI (Ai2)32.2B≈21.0 GB at 4-bit
- FLUX.2 [dev]Black Forest Labs32.2B≈20.9 GB at 4-bit1 also selling it hosted
- sarvam-30bSarvam AI32.2B≈20.9 GB at 4-bit
- Baichuan M2 32Bbaichuan32.0B≈20.8 GB at 4-bit2 also selling it hosted
- DeepSeek R1 Distill QWEN 32BDeepSeek32.0B≈20.8 GB at 4-bit6 also selling it hosted
- EXAONE-Deep-32BLGAI-EXAONE32.0B≈20.8 GB at 4-bit
- GLM-4-32B-0414Z.ai32.0B≈20.8 GB at 4-bit3 also selling it hosted
- Olmo-2-0325-32B-InstructAllen Institute for AI (Ai2)32.0B≈20.8 GB at 4-bit
- Olmo-3.1-32B-InstructAllen Institute for AI (Ai2)32.0B≈20.8 GB at 4-bit1 also selling it hosted
- OlympicCoder-32Bopen-r132.0B≈20.8 GB at 4-bit
- OpenReasoning-Nemotron-32BNVIDIA32.0B≈20.8 GB at 4-bit1 also selling it hosted
- OpenThinker-32Bopen-thoughts32.0B≈20.8 GB at 4-bit
- QwQ-32BAlibaba32.0B≈20.8 GB at 4-bit7 also selling it hosted
- QwQ-32B-PreviewAlibaba32.0B≈20.8 GB at 4-bit2 also selling it hosted
- QwQ-R1984-32BVIDraft32.0B≈20.8 GB at 4-bit
- Qwen-SEA-LION-v4-32B-ITaisingapore32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen2.5-32BAlibaba32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen2.5-VL-32B-InstructAlibaba32.0B≈20.8 GB at 4-bit3 also selling it hosted
- QwenLong-L1-32BTongyi-Zhiwen32.0B≈20.8 GB at 4-bit
- Sky-T1-32B-PreviewNovaSky-AI32.0B≈20.8 GB at 4-bit1 also selling it hosted
- UIGEN-X-32B-0727Tesslate32.0B≈20.8 GB at 4-bit
- s1-32Bsimplescaling32.0B≈20.8 GB at 4-bit
- xLAM-2-32b-fc-rSalesforce32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen3-Omni-30B-A3B-CaptionerAlibaba31.7B≈20.6 GB at 4-bit
- NVIDIA Nemotron 3.5 Lightning 30B A3BNVIDIA31.6B≈20.5 GB at 4-bit2 also selling it hosted
- Nemotron 3 Nano 30B A3BNVIDIA31.6B≈20.5 GB at 4-bit10 also selling it hosted
- Nemotron-Cascade-2-30B-A3BNVIDIA31.6B≈20.5 GB at 4-bit
- phonellm-alpha-1Pipecat31.6B≈20.5 GB at 4-bit
- Gemma 4 31BGoogle31.3B≈20.3 GB at 4-bit15 also selling it hosted
- Qwen3 VL 30B A3B InstructAlibaba31.1B≈20.2 GB at 4-bit8 also selling it hosted
- Qwen3 VL 30B A3B ThinkingAlibaba31.1B≈20.2 GB at 4-bit7 also selling it hosted
- Gemma-4-31B-it-assistantGoogle31.0B≈20.2 GB at 4-bit
- Gemma-4-31B-it-pearlpearl-ai31.0B≈20.2 GB at 4-bit
- Qwen3-Coder-30B-A3B-Instruct-FP8Alibaba30.5B≈19.8 GB at 4-bit
- MiroThinker-v1.5-30BMiroMind AI30.5B≈19.8 GB at 4-bit
- Qwen3 30B A3BAlibaba30.5B≈19.8 GB at 4-bit11 also selling it hosted
- Qwen3 30B A3B Instruct 2507Alibaba30.5B≈19.8 GB at 4-bit8 also selling it hosted
- Qwen3 30B A3B Thinking 2507Alibaba30.5B≈19.8 GB at 4-bit4 also selling it hosted
- Qwen3 Coder 30B A3B InstructAlibaba30.5B≈19.8 GB at 4-bit11 also selling it hosted
- Tongyi-DeepResearch-30B-A3BAlibaba-NLP30.5B≈19.8 GB at 4-bit
- Hy-MT2-30B-A3BTencent Hunyuan30.0B≈19.5 GB at 4-bit1 also selling it hosted
- Nemotron-Labs-Audex-30B-A3BNVIDIA30.0B≈19.5 GB at 4-bit
- Ovis2.6-30B-A3BATH-MaaS30.0B≈19.5 GB at 4-bit
- Qwen3 Omni 30B A3B InstructAlibaba30.0B≈19.5 GB at 4-bit2 also selling it hosted
- Qwen3 Omni 30B A3B ThinkingAlibaba30.0B≈19.5 GB at 4-bit2 also selling it hosted
- QwenLong-L1.5-30B-A3BTongyi-Zhiwen30.0B≈19.5 GB at 4-bit
- TildeOpen-30bTildeAI30.0B≈19.5 GB at 4-bit
- Wizard-Vicuna-30B-UncensoredQuixi AI30.0B≈19.5 GB at 4-bit
- WizardLM-30B-UncensoredQuixi AI30.0B≈19.5 GB at 4-bit
- granite-4.1-30bIBM Granite30.0B≈19.5 GB at 4-bit
- Muse Glimmer 30BMeta29.8B≈19.4 GB at 4-bit8 also selling it hosted
- stepvideo-t2vStepFun29.3B≈19.1 GB at 4-bit
- medgemma-27b-itGoogle28.8B≈18.7 GB at 4-bit
- translategemma-27b-itGoogle28.8B≈18.7 GB at 4-bit
- ERNIE 4.5 VL 28B A3BBaidu28.0B≈18.2 GB at 4-bit2 also selling it hosted
- ERNIE-4.5-VL-28B-A3B-ThinkingBaidu28.0B≈18.2 GB at 4-bit1 also selling it hosted
- Huihui-Qwen3.8-27B-abliteratedhuihui-ai27.8B≈18.1 GB at 4-bit
- Qwen3.5-27BAlibaba27.8B≈18.1 GB at 4-bit8 also selling it hosted
- Qwen3.5-27B-Claude-4.6-Opus-Reasoning-DistilledJackrong27.8B≈18.1 GB at 4-bit
- Qwen3.6 27BAlibaba27.8B≈18.1 GB at 4-bit12 also selling it hosted
- Qwen3.8 27BAlibaba27.8B≈18.1 GB at 4-bit12 also selling it hosted