Conversation models that run on 96 GB
600 models with published weights that fit in 96 GB — a MacBook Pro with 96. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 102 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 600 in all, a hundred to a page; this is page 5 of 6.
An Apple configuration for people who intend to run models rather than occasionally try one. Comfortably holds the seventy-billion class with a very long context, or a mixture-of-experts model whose active parameters are few but whose weights are all resident.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- llama2_7b_chat_uncensoredgeorgesung6.7B≈4.4 GB at 4-bit
- LlamaGuard-7bMeta Llama6.7B≈4.4 GB at 4-bit1 also selling it hosted
- YuE-s1-7B-anneal-en-cotMultimodal Art Projection6.2B≈4.0 GB at 4-bit
- Yi-6B01-ai6.1B≈3.9 GB at 4-bit1 also selling it hosted
- Yi-6B-200K01-ai6.0B≈3.9 GB at 4-bit
- gpt-j-6bEleutherAI6.0B≈3.9 GB at 4-bit
- pygmalion-6bPygmalion6.0B≈3.9 GB at 4-bit
- DeciLM-6bDeci AI5.7B≈3.7 GB at 4-bit
- Mage-VLMicrosoft4.7B≈3.1 GB at 4-bit
- Qwen-Drive-1.0-4BAlibaba4.5B≈3.0 GB at 4-bit
- Llama-3.1-Minitron-4B-Width-BaseNVIDIA4.5B≈2.9 GB at 4-bit
- medgemma-1.5-4b-itGoogle4.3B≈2.8 GB at 4-bit
- medgemma-4b-itGoogle4.3B≈2.8 GB at 4-bit
- NeoHorse-1-4BTokenRhythm4.2B≈2.7 GB at 4-bit
- Nanbeige4.2-3BNanbeige LLM Lab4.2B≈2.7 GB at 4-bit
- Phi-3-vision-128k-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Phi-3.5-vision-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Spark-X2.5-4BSparkLLM4.1B≈2.7 GB at 4-bit
- MiniCPM-V-4OpenBMB4.1B≈2.6 GB at 4-bit
- AgentCPM-ExploreOpenBMB4.0B≈2.6 GB at 4-bit
- Jan-nano-128kMenlo Research4.0B≈2.6 GB at 4-bit
- Jan-v1-4BJan4.0B≈2.6 GB at 4-bit
- Mellum-4b-baseJetBrains4.0B≈2.6 GB at 4-bit
- GELab-Zero-4B-previewstepfun-ai4.0B≈2.6 GB at 4-bit
- Gemma-3-Gaia-PT-BR-4b-itCEIA-UFG4.0B≈2.6 GB at 4-bit
- LocoTrainer-4BLocoreMind4.0B≈2.6 GB at 4-bit
- MiniCPM3-4BOpenBMB4.0B≈2.6 GB at 4-bit
- Nemotron-Mini-4B-InstructNVIDIA4.0B≈2.6 GB at 4-bit
- Qwen3-4b-Z-Image-Engineer-V4BennyDaBall4.0B≈2.6 GB at 4-bit
- Qwen3-VL-4B-InstructAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- R-4BYannQi4.0B≈2.6 GB at 4-bit
- Youtu-VL-4B-InstructTencent Hunyuan4.0B≈2.6 GB at 4-bit
- gemma-3-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- medgemma-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- t5gemma-2-4b-4bGoogle4.0B≈2.6 GB at 4-bit
- Nanbeige4.1-3BNanbeige LLM Lab3.9B≈2.6 GB at 4-bit
- Phi-4-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Phi-3-mini-128k-instructMicrosoft3.8B≈2.5 GB at 4-bit2 also selling it hosted
- Phi-3-mini-4k-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Phi-3.5-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Hy-Embodied-0.5Tencent Hunyuan3.8B≈2.5 GB at 4-bit
- blip2-opt-2.7bSalesforce3.7B≈2.4 GB at 4-bit
- HyperCLOVAX-SEED-Vision-Instruct-3BHyperCLOVA X3.7B≈2.4 GB at 4-bit
- deepseek-vl2-tinydeepseek-ai3.4B≈2.2 GB at 4-bit
- nllb-200-3.3BAI at Meta3.3B≈2.1 GB at 4-bit
- Llama 3.2 3B InstructMeta3.2B≈2.1 GB at 4-bit14 also selling it hosted
- llama-3.2-Korean-Bllossom-3BBllossom3.2B≈2.1 GB at 4-bit
- imp-v1-3bMILVLG3.2B≈2.1 GB at 4-bit
- LFM2.5-VL-3BLiquid AI3.1B≈2.0 GB at 4-bit
- Qwen2.5-3BAlibaba3.1B≈2.0 GB at 4-bit
- Qwen2.5-3B-InstructAlibaba3.1B≈2.0 GB at 4-bit
- VibeThinker-3BWeiboAI3.1B≈2.0 GB at 4-bit
- SmolLM3-3BHugging Face Smol Models Research3.1B≈2.0 GB at 4-bit
- OpenELM-3B-InstructApple3.0B≈2.0 GB at 4-bit
- SmolLM3-3B-BaseHuggingFaceTB3.0B≈2.0 GB at 4-bit
- bitnet_b1_58-3B1bitLLM3.0B≈2.0 GB at 4-bit
- open_llama_3bOpenLM Research3.0B≈2.0 GB at 4-bit
- open_llama_3b_v2OpenLM Research3.0B≈2.0 GB at 4-bit
- orca_mini_3bpankajmathur3.0B≈2.0 GB at 4-bit
- paligemma2-3b-pt-224Google3.0B≈2.0 GB at 4-bit
- proxy-lite-3bconvergence-ai3.0B≈2.0 GB at 4-bit
- stablelm-3b-4e1tStability AI3.0B≈2.0 GB at 4-bit
- paligemma-3b-pt-224Google2.9B≈1.9 GB at 4-bit
- stablelm-zephyr-3bStability AI2.8B≈1.8 GB at 4-bit
- dolphin-2_6-phi-2Dolphin2.8B≈1.8 GB at 4-bit
- phi-2Microsoft2.8B≈1.8 GB at 4-bit
- Solidity-LLMChainGPT2.8B≈1.8 GB at 4-bit
- gpt-neo-2.7BEleutherAI2.7B≈1.8 GB at 4-bit
- gemma-2-2bGoogle2.6B≈1.7 GB at 4-bit
- gemma-2-2b-itGoogle2.6B≈1.7 GB at 4-bit
- gemma-2-2b-jpn-itGoogle2.6B≈1.7 GB at 4-bit
- LFM2-2.6B-TranscriptLiquid AI2.6B≈1.7 GB at 4-bit
- Ovis2.5-2BATH-MaaS2.6B≈1.7 GB at 4-bit
- LFM2-2.6BLiquid AI2.6B≈1.7 GB at 4-bit
- LFM2-2.6B-ExpLiquid AI2.6B≈1.7 GB at 4-bit
- MiniCPM5-2BOpenBMB2.5B≈1.6 GB at 4-bit
- Octopus-v2Nexa AI2.5B≈1.6 GB at 4-bit
- gemma-2bGoogle2.5B≈1.6 GB at 4-bit
- gemma-2b-itGoogle2.5B≈1.6 GB at 4-bit1 also selling it hosted
- Cosmos-Reason2-2BNVIDIA2.4B≈1.6 GB at 4-bit
- EXAONE-3.5-2.4B-InstructLGAI-EXAONE2.4B≈1.6 GB at 4-bit
- SmolVLM2-2.2B-InstructHugging Face Smol Models Research2.2B≈1.5 GB at 4-bit
- Marlin-2BNemo Station2.2B≈1.4 GB at 4-bit
- Qwen2-VL-2B-InstructAlibaba2.2B≈1.4 GB at 4-bit1 also selling it hosted
- Qwen3-VL-2B-InstructAlibaba2.1B≈1.4 GB at 4-bit
- MiniCPM-2B-sft-fp32OpenBMB2.0B≈1.3 GB at 4-bit
- Qwen3.5-2BAlibaba2.0B≈1.3 GB at 4-bit2 also selling it hosted
- gemma-1.1-2b-itGoogle2.0B≈1.3 GB at 4-bit
- helium-1-preview-2bKyutai2.0B≈1.3 GB at 4-bit
- Youtu-LLM-2BTencent Hunyuan2.0B≈1.3 GB at 4-bit
- moondream2vikhyatk1.9B≈1.3 GB at 4-bit
- VibeThinker-1.5BWeiboAI1.8B≈1.2 GB at 4-bit
- Qwen3.6-27B-DFlashZ Lab1.7B≈1.1 GB at 4-bit
- SmolLM-1.7BHuggingFaceTB1.7B≈1.1 GB at 4-bit
- SmolLM2-1.7BHuggingFaceTB1.7B≈1.1 GB at 4-bit
- stablelm-2-1_6bStability AI1.6B≈1.1 GB at 4-bit
- stablelm-2-zephyr-1_6bStability AI1.6B≈1.1 GB at 4-bit
- LFM2.5-VL-1.6BLiquid AI1.6B≈1.0 GB at 4-bit
- LFM2-VL-1.6BLiquid AI1.6B≈1.0 GB at 4-bit
- Qwen2.5-1.5BAlibaba1.5B≈1.0 GB at 4-bit