Models that run on 96 GB
1115 models with published weights that fit in 96 GB — a MacBook Pro with 96. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 102 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 1,115 in all, a hundred to a page; this is page 7 of 12.
An Apple configuration for people who intend to run models rather than occasionally try one. Comfortably holds the seventy-billion class with a very long context, or a mixture-of-experts model whose active parameters are few but whose weights are all resident.
- Qwen2.5-Coder-7B-InstructAlibaba7.0B≈4.5 GB at 4-bit4 also selling it hosted
- Seed-X-PPO-7BByteDance Seed7.0B≈4.5 GB at 4-bit
- TAIDE-LX-7B-Chattaide7.0B≈4.5 GB at 4-bit
- UI-TARS-7B-SFTByteDance Seed7.0B≈4.5 GB at 4-bit
- Vistral-7B-ChatViet-Mistral7.0B≈4.5 GB at 4-bit
- WizardLM-7B-UncensoredQuixi AI7.0B≈4.5 GB at 4-bit
- chinese-alpaca-2-7bJoint Laboratory of HIT and iFLYTEK Research (HFL)7.0B≈4.5 GB at 4-bit
- codegemma-7b-itGoogle7.0B≈4.5 GB at 4-bit1 also selling it hosted
- deepseek-coder-7b-instruct-v1.5deepseek-ai7.0B≈4.5 GB at 4-bit1 also selling it hosted
- deepseek-llm-7b-basedeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-llm-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-math-7b-instructdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-vl-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- e5-mistral-7b-instructintfloat7.0B≈4.5 GB at 4-bit
- gemma-1.1-7b-itGoogle7.0B≈4.5 GB at 4-bit1 also selling it hosted
- gte-Qwen2-7B-instructAlibaba-NLP7.0B≈4.5 GB at 4-bit
- internlm-xcomposer2d5-7binternlm7.0B≈4.5 GB at 4-bit
- llava-v1.6-mistral-7b-hfLlava Hugging Face7.0B≈4.5 GB at 4-bit
- mixtral-7b-8expertDiscoResearch7.0B≈4.5 GB at 4-bit
- olmOCR-2-7B-1025Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit1 also selling it hosted
- olmOCR-7B-0725-FP8Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit1 also selling it hosted
- olmOCR-7B-0825Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit1 also selling it hosted
- open-calm-7bCyberAgent7.0B≈4.5 GB at 4-bit
- rwkv-4-pile-7bBlinkDL7.0B≈4.5 GB at 4-bit
- speed-embedding-7b-instructHaon-Chen7.0B≈4.5 GB at 4-bit
- stablelm-base-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- stablelm-tuned-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- vicuna-7b-v1.5Large Model Systems Organization7.0B≈4.5 GB at 4-bit
- xgen-7b-8k-baseSalesforce7.0B≈4.5 GB at 4-bit
- granite-4.0-h-tinyIBM Granite6.9B≈4.5 GB at 4-bit
- NuminaMath-7B-TIRProject-Numina6.9B≈4.5 GB at 4-bit
- OLMo-7BAi26.9B≈4.5 GB at 4-bit
- Magicoder-S-DS-6.7BIntellligent Software Engineering (iSE)6.7B≈4.4 GB at 4-bit
- deepseek-coder-6.7b-instructDeepSeek6.7B≈4.4 GB at 4-bit
- meditron-7bEPFL LLM Team6.7B≈4.4 GB at 4-bit
- CodeLlama-7b-Instruct-hfCode Llama6.7B≈4.4 GB at 4-bit
- CodeLlama-7b-hfCode Llama6.7B≈4.4 GB at 4-bit
- sqlcoder-7b-2Defog.ai6.7B≈4.4 GB at 4-bit
- Llama-2-7b-chat-hfMeta Llama6.7B≈4.4 GB at 4-bit
- Llama-2-7b-hfMeta Llama6.7B≈4.4 GB at 4-bit
- llama-7bhuggyllama6.7B≈4.4 GB at 4-bit
- llama2_7b_chat_uncensoredgeorgesung6.7B≈4.4 GB at 4-bit
- LlamaGuard-7bMeta Llama6.7B≈4.4 GB at 4-bit1 also selling it hosted
- dinov3-vit7b16-pretrain-lvd1689mfacebook6.7B≈4.4 GB at 4-bit
- YuE-s1-7B-anneal-en-cotMultimodal Art Projection6.2B≈4.0 GB at 4-bit
- Z-ImageTongyi-MAI6.2B≈4.0 GB at 4-bit
- Z-Image-TurboTongyi-MAI6.2B≈4.0 GB at 4-bit
- Yi-6B01-ai6.1B≈3.9 GB at 4-bit1 also selling it hosted
- Yi-6B-200K01-ai6.0B≈3.9 GB at 4-bit
- gpt-j-6bEleutherAI6.0B≈3.9 GB at 4-bit
- pygmalion-6bPygmalion6.0B≈3.9 GB at 4-bit
- Chroma-4BFlashLabs5.9B≈3.8 GB at 4-bit
- higgs-tts-2-3b-baseBoson AI5.8B≈3.8 GB at 4-bit
- DeciLM-6bDeci AI5.7B≈3.7 GB at 4-bit
- CogVideoX-5bZ.ai5.6B≈3.6 GB at 4-bit
- Qwen2.5-Omni-3BAlibaba5.5B≈3.6 GB at 4-bit
- chandra-ocr-2Datalab5.3B≈3.4 GB at 4-bit
- Gemma-4-E2B-itGoogle5.1B≈3.3 GB at 4-bit1 also selling it hosted
- CogVideoX-5b-I2Vzai-org5.0B≈3.2 GB at 4-bit
- CogVideoX1.5-5B-SATzai-org5.0B≈3.2 GB at 4-bit
- Wan2.2-TI2V-5BWan-AI5.0B≈3.2 GB at 4-bit
- Wan2.2-TI2V-5B-DiffusersWan-AI5.0B≈3.2 GB at 4-bit
- acestep-v15-xl-turboACE-Step5.0B≈3.2 GB at 4-bit
- translategemma-4b-itGoogle5.0B≈3.2 GB at 4-bit
- Mage-VLMicrosoft4.7B≈3.1 GB at 4-bit
- higgs-tts-3-4bBoson AI4.7B≈3.0 GB at 4-bit
- s2-proFish Audio4.6B≈3.0 GB at 4-bit1 also selling it hosted
- NuExtract3NuMind4.5B≈3.0 GB at 4-bit
- Qwen-Drive-1.0-4BAlibaba4.5B≈3.0 GB at 4-bit
- Llama-3.1-Minitron-4B-Width-BaseNVIDIA4.5B≈2.9 GB at 4-bit
- Voxtral-Mini-4B-Realtime-2602Mistral AI4.4B≈2.9 GB at 4-bit
- IF-I-XL-v1.0DeepFloyd4.3B≈2.8 GB at 4-bit
- Gemma 3 4BGoogle4.3B≈2.8 GB at 4-bit6 also selling it hosted
- medgemma-1.5-4b-itGoogle4.3B≈2.8 GB at 4-bit
- medgemma-4b-itGoogle4.3B≈2.8 GB at 4-bit
- NeoHorse-1-4BTokenRhythm4.2B≈2.7 GB at 4-bit
- Nanbeige4.2-3BNanbeige LLM Lab4.2B≈2.7 GB at 4-bit
- Phi-3-vision-128k-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Phi-3.5-vision-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Spark-X2.5-4BSparkLLM4.1B≈2.7 GB at 4-bit
- MiniCPM-V-4OpenBMB4.1B≈2.6 GB at 4-bit
- AgentCPM-ExploreOpenBMB4.0B≈2.6 GB at 4-bit
- DASD-4B-ThinkingAlibaba Cloud Apsara Lab4.0B≈2.6 GB at 4-bit
- Jan-nanoMenlo Research4.0B≈2.6 GB at 4-bit
- Jan-nano-128kMenlo Research4.0B≈2.6 GB at 4-bit
- Jan-v1-4BJan4.0B≈2.6 GB at 4-bit
- fable-tracesAliesTaha4.0B≈2.6 GB at 4-bit
- Mellum-4b-baseJetBrains4.0B≈2.6 GB at 4-bit
- Llasa-3BHKUST Audio4.0B≈2.6 GB at 4-bit
- F2LLM-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- F2LLM-v2-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- FLUX.2-klein-base-4BBlack Forest Labs4.0B≈2.6 GB at 4-bit
- GELab-Zero-4B-previewstepfun-ai4.0B≈2.6 GB at 4-bit
- Gemma-3-Gaia-PT-BR-4b-itCEIA-UFG4.0B≈2.6 GB at 4-bit
- LocoOperator-4BLocoreMind4.0B≈2.6 GB at 4-bit
- LocoTrainer-4BLocoreMind4.0B≈2.6 GB at 4-bit
- MiniCPM3-4BOpenBMB4.0B≈2.6 GB at 4-bit
- Nemotron-Mini-4B-InstructNVIDIA4.0B≈2.6 GB at 4-bit
- OmniNeural-4BNexa AI4.0B≈2.6 GB at 4-bit
- Qwen3 4BAlibaba4.0B≈2.6 GB at 4-bit3 also selling it hosted