Models that run on 96 GB
1115 models with published weights that fit in 96 GB — a MacBook Pro with 96. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 102 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 1,115 in all, a hundred to a page; this is page 1 of 12.
An Apple configuration for people who intend to run models rather than occasionally try one. Comfortably holds the seventy-billion class with a very long context, or a mixture-of-experts model whose active parameters are few but whose weights are all resident.
- Llama-3.2-90B-Vision-Instruct90.0B≈58.5 GB at 4-bit7 also selling it hosted
- HunyuanImage-3.0Tencent Hunyuan83.0B≈54.0 GB at 4-bit
- HunyuanImage-3.0-InstructTencent Hunyuan83.0B≈54.0 GB at 4-bit
- Qwen3 Next 80B A3B InstructAlibaba81.3B≈52.9 GB at 4-bit12 also selling it hosted
- Qwen3 Next 80B A3B ThinkingAlibaba81.3B≈52.9 GB at 4-bit9 also selling it hosted
- idefics-80b-instructHuggingFaceM480.0B≈52.0 GB at 4-bit
- Qwen3 Coder NextAlibaba79.7B≈51.8 GB at 4-bit7 also selling it hosted
- NVLM-D-72BNVIDIA79.4B≈51.6 GB at 4-bit
- InternVL3-78BOpenGVLab78.4B≈51.0 GB at 4-bit1 also selling it hosted
- calme-3.2-instruct-78bMaziyarPanahi78.0B≈50.7 GB at 4-bit
- LongCat-NextMeituan LongCat74.3B≈48.3 GB at 4-bit
- Qwen2.5 VL 72B InstructAlibaba73.4B≈47.7 GB at 4-bit7 also selling it hosted
- Kimi-Dev-72BMoonshot AI72.7B≈47.3 GB at 4-bit
- Magnum v4 72BAnthracite72.7B≈47.3 GB at 4-bit1 also selling it hosted
- Qwen2-72BAlibaba72.7B≈47.3 GB at 4-bit
- Qwen2.5 72B InstructAlibaba72.7B≈47.3 GB at 4-bit12 also selling it hosted
- Virtuoso LargeArcee AI72.7B≈47.3 GB at 4-bit
- Smaug-72B-v0.1Abacus.AI, Inc.72.3B≈47.0 GB at 4-bit
- Qwen-72BAlibaba72.3B≈47.0 GB at 4-bit
- Qwen1.5-72B-ChatAlibaba72.3B≈47.0 GB at 4-bit1 also selling it hosted
- KAT-Dev-72B-ExpKwaipilot72.0B≈46.8 GB at 4-bit1 also selling it hosted
- Molmo-72B-0924Allen Institute for AI (Ai2)72.0B≈46.8 GB at 4-bit
- QVQ-72B-PreviewAlibaba72.0B≈46.8 GB at 4-bit1 also selling it hosted
- Qwen 2 VL 72B InstructAlibaba72.0B≈46.8 GB at 4-bit4 also selling it hosted
- Qwen-72B-ChatAlibaba72.0B≈46.8 GB at 4-bit
- Qwen2-72B-InstructAlibaba72.0B≈46.8 GB at 4-bit3 also selling it hosted
- UI-TARS-72B-DPOByteDance Seed72.0B≈46.8 GB at 4-bit
- dolphin-2.9.2-qwen2-72bdphn72.0B≈46.8 GB at 4-bit1 also selling it hosted
- magnum-v1-72banthracite-org72.0B≈46.8 GB at 4-bit
- Llama3-ChatQA-1.5-70BNVIDIA70.6B≈45.9 GB at 4-bit
- Athene-70BNexusflow70.6B≈45.9 GB at 4-bit
- Hermes 3 70B InstructNous Research70.6B≈45.9 GB at 4-bit4 also selling it hosted
- Hermes 4 70BNous Research70.6B≈45.9 GB at 4-bit4 also selling it hosted
- Higgs-Llama-3-70BBoson AI70.6B≈45.9 GB at 4-bit
- Llama 3.1 Instruct (70B)Meta70.6B≈45.9 GB at 4-bit7 also selling it hosted
- Llama 3.3 70B InstructMeta70.6B≈45.9 GB at 4-bit29 also selling it hosted
- Llama-3.1-Nemotron-70B-Instruct-HFNVIDIA70.6B≈45.9 GB at 4-bit1 also selling it hosted
- R1 Distill Llama 70BDeepSeek70.6B≈45.9 GB at 4-bit12 also selling it hosted
- Apertus-70B-2509swiss-ai70.0B≈45.5 GB at 4-bit
- Apertus-70B-Instruct-2509swiss-ai70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Apertus-v1.5-70Bswiss-ai70.0B≈45.5 GB at 4-bit1 also selling it hosted
- CodeLlama-70b-hfCode Llama70.0B≈45.5 GB at 4-bit
- L3 70B Euryale V2.1Sao10K70.0B≈45.5 GB at 4-bit3 also selling it hosted
- L3.3-70B-Euryale-v2.3Sao10K70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Llama-2-70bMeta Llama70.0B≈45.5 GB at 4-bit
- Llama-2-70b-chatMeta Llama70.0B≈45.5 GB at 4-bit2 also selling it hosted
- Llama-2-70b-chat-hfMeta70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Llama-3-Groq-70B-Tool-UseGroq70.0B≈45.5 GB at 4-bit
- Llama-3.1-Nemotron-70B-InstructNVIDIA70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Llama-xLAM-2-70b-fc-rSalesforce70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Llama3-OpenBioLLM-70Baaditya70.0B≈45.5 GB at 4-bit
- Meta-Llama-3-70B-InstructMeta70.0B≈45.5 GB at 4-bit6 also selling it hosted
- Meta-Llama-3.1-70B-InstructMeta70.0B≈45.5 GB at 4-bit7 also selling it hosted
- Platypus2-70B-instructgarage-bAInd70.0B≈45.5 GB at 4-bit
- Reflection-Llama-3.1-70Bmattshumer70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Smaug-Llama-3-70B-InstructAbacus.AI, Inc.70.0B≈45.5 GB at 4-bit
- WizardLM-70B-V1.0WizardLM Team70.0B≈45.5 GB at 4-bit
- Xwin-LM-70B-V0.1Xwin-LM70.0B≈45.5 GB at 4-bit
- dolphin-2.9.1-llama-3-70bcognitivecomputations70.0B≈45.5 GB at 4-bit1 also selling it hosted
- llama-3-70B-Instruct-abliteratedfailspy70.0B≈45.5 GB at 4-bit
- lzlv_70b_fp16_hflizpreciatior70.0B≈45.5 GB at 4-bit1 also selling it hosted
- med42-70bM42 Health70.0B≈45.5 GB at 4-bit
- meditron-70bEPFL LLM Team70.0B≈45.5 GB at 4-bit
- tulu-2-dpo-70bAllen Institute for AI (Ai2)70.0B≈45.5 GB at 4-bit
- LongCat-Flash-LiteMeituan LongCat69.1B≈44.9 GB at 4-bit
- CodeLlama-70b-Instruct-hfCode Llama69.0B≈44.8 GB at 4-bit
- sqlcoder-70b-alphadefog69.0B≈44.8 GB at 4-bit
- Llama-2-70b-hfMeta Llama69.0B≈44.8 GB at 4-bit
- NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4NVIDIA67.2B≈43.7 GB at 4-bit
- deepseek-llm-67b-chatdeepseek-ai67.0B≈43.6 GB at 4-bit
- opt-66bfacebook66.0B≈42.9 GB at 4-bit
- Jamba-v0.1AI21 Labs51.6B≈33.5 GB at 4-bit
- Llama-3_1-Nemotron-51B-InstructNVIDIA51.5B≈33.5 GB at 4-bit
- Kimi-Linear-48B-A3B-InstructMoonshot AI49.1B≈31.9 GB at 4-bit
- Llama-3_3-Nemotron-Super-49B-v1NVIDIA49.0B≈31.9 GB at 4-bit
- dolphin-2.5-mixtral-8x7bDolphin46.7B≈30.4 GB at 4-bit
- GRIN-MoEMicrosoft41.9B≈27.2 GB at 4-bit
- Phi-3.5-MoE-instructMicrosoft41.9B≈27.2 GB at 4-bit1 also selling it hosted
- falcon-40bTechnology Innovation Institute41.8B≈27.2 GB at 4-bit
- IQuest-Coder-V1-40B-InstructIQuest40.0B≈26.0 GB at 4-bit
- falcon-40b-instructTechnology Innovation Institute40.0B≈26.0 GB at 4-bit
- IQuest-Coder-V1-40B-Loop-InstructIQuest39.8B≈25.9 GB at 4-bit
- Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingDavidAU39.5B≈25.7 GB at 4-bit
- CLIP-ViT-bigG-14-laion2B-39B-b160kLAION eV39.0B≈25.4 GB at 4-bit
- Seed-OSS-36B-InstructByteDance Seed36.2B≈23.5 GB at 4-bit1 also selling it hosted
- K2-Horizon-MoVA-36B-A4BInstitute of Foundation Models36.0B≈23.4 GB at 4-bit
- Skyfall 36B V2TheDrummer36.0B≈23.4 GB at 4-bit1 also selling it hosted
- BigBang-v1The Endless Frontier36.0B≈23.4 GB at 4-bit
- Huihui-Qwen3.5-35B-A3B-abliteratedhuihui-ai36.0B≈23.4 GB at 4-bit
- Ornith-1.5-35B-A3BOrnith36.0B≈23.4 GB at 4-bit
- Qwable-v1lordx6436.0B≈23.4 GB at 4-bit
- Qwen3.5-35B-A3BAlibaba36.0B≈23.4 GB at 4-bit9 also selling it hosted
- Qwen3.6 35B A3BAlibaba36.0B≈23.4 GB at 4-bit10 also selling it hosted
- Agents-A1Intern Science35.1B≈22.8 GB at 4-bit
- Nex-N2-MiniNEX AGI35.1B≈22.8 GB at 4-bit
- Nex-N2.5-miniNEX AGI35.1B≈22.8 GB at 4-bit1 also selling it hosted
- Thomson-1.0-SmallThomson Reuters35.1B≈22.8 GB at 4-bit
- XYZ-Aquila-miniXYZAILab35.1B≈22.8 GB at 4-bit
- Holo3-35B-A3BH Company35.0B≈22.8 GB at 4-bit1 also selling it hosted
- Ornith-1.0-35Bdeepreinforce-ai35.0B≈22.8 GB at 4-bit1 also selling it hosted