Models that run on 256 GB
1161 models with published weights that fit in 256 GB — a Mac Studio with 256. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 274 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 1,161 in all, a hundred to a page; this is page 8 of 12.
A Mac Studio at the top of its configuration, or a serious machine. Almost everything openly published fits, including the large mixtures of experts. At this point the constraint is no longer whether the model loads but whether it generates fast enough to be worth waiting for.
- CogVideoX-5bZ.ai5.6B≈3.6 GB at 4-bit
- Qwen2.5-Omni-3BAlibaba5.5B≈3.6 GB at 4-bit
- chandra-ocr-2Datalab5.3B≈3.4 GB at 4-bit
- Gemma-4-E2B-itGoogle5.1B≈3.3 GB at 4-bit1 also selling it hosted
- CogVideoX-5b-I2Vzai-org5.0B≈3.2 GB at 4-bit
- CogVideoX1.5-5B-SATzai-org5.0B≈3.2 GB at 4-bit
- Wan2.2-TI2V-5BWan-AI5.0B≈3.2 GB at 4-bit
- Wan2.2-TI2V-5B-DiffusersWan-AI5.0B≈3.2 GB at 4-bit
- acestep-v15-xl-turboACE-Step5.0B≈3.2 GB at 4-bit
- translategemma-4b-itGoogle5.0B≈3.2 GB at 4-bit
- Mage-VLMicrosoft4.7B≈3.1 GB at 4-bit
- higgs-tts-3-4bBoson AI4.7B≈3.0 GB at 4-bit
- s2-proFish Audio4.6B≈3.0 GB at 4-bit1 also selling it hosted
- NuExtract3NuMind4.5B≈3.0 GB at 4-bit
- Qwen-Drive-1.0-4BAlibaba4.5B≈3.0 GB at 4-bit
- Llama-3.1-Minitron-4B-Width-BaseNVIDIA4.5B≈2.9 GB at 4-bit
- Voxtral-Mini-4B-Realtime-2602Mistral AI4.4B≈2.9 GB at 4-bit
- IF-I-XL-v1.0DeepFloyd4.3B≈2.8 GB at 4-bit
- Gemma 3 4BGoogle4.3B≈2.8 GB at 4-bit6 also selling it hosted
- medgemma-1.5-4b-itGoogle4.3B≈2.8 GB at 4-bit
- medgemma-4b-itGoogle4.3B≈2.8 GB at 4-bit
- NeoHorse-1-4BTokenRhythm4.2B≈2.7 GB at 4-bit
- Nanbeige4.2-3BNanbeige LLM Lab4.2B≈2.7 GB at 4-bit
- Phi-3-vision-128k-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Phi-3.5-vision-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Spark-X2.5-4BSparkLLM4.1B≈2.7 GB at 4-bit
- MiniCPM-V-4OpenBMB4.1B≈2.6 GB at 4-bit
- AgentCPM-ExploreOpenBMB4.0B≈2.6 GB at 4-bit
- DASD-4B-ThinkingAlibaba Cloud Apsara Lab4.0B≈2.6 GB at 4-bit
- Jan-nanoMenlo Research4.0B≈2.6 GB at 4-bit
- Jan-nano-128kMenlo Research4.0B≈2.6 GB at 4-bit
- Jan-v1-4BJan4.0B≈2.6 GB at 4-bit
- fable-tracesAliesTaha4.0B≈2.6 GB at 4-bit
- Mellum-4b-baseJetBrains4.0B≈2.6 GB at 4-bit
- Llasa-3BHKUST Audio4.0B≈2.6 GB at 4-bit
- F2LLM-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- F2LLM-v2-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- FLUX.2-klein-base-4BBlack Forest Labs4.0B≈2.6 GB at 4-bit
- GELab-Zero-4B-previewstepfun-ai4.0B≈2.6 GB at 4-bit
- Gemma-3-Gaia-PT-BR-4b-itCEIA-UFG4.0B≈2.6 GB at 4-bit
- LocoOperator-4BLocoreMind4.0B≈2.6 GB at 4-bit
- LocoTrainer-4BLocoreMind4.0B≈2.6 GB at 4-bit
- MiniCPM3-4BOpenBMB4.0B≈2.6 GB at 4-bit
- Nemotron-Mini-4B-InstructNVIDIA4.0B≈2.6 GB at 4-bit
- OmniNeural-4BNexa AI4.0B≈2.6 GB at 4-bit
- Qwen3 4BAlibaba4.0B≈2.6 GB at 4-bit3 also selling it hosted
- Qwen3-4B-Instruct-2507Alibaba4.0B≈2.6 GB at 4-bit2 also selling it hosted
- Qwen3-4B-Thinking-2507Alibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3-4b-Z-Image-Engineer-V4BennyDaBall4.0B≈2.6 GB at 4-bit
- Qwen3-Embedding-4BAlibaba4.0B≈2.6 GB at 4-bit4 also selling it hosted
- Qwen3-Reranker-4BAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3-VL-4B-InstructAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3-VL-4B-ThinkingAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3.5-4BAlibaba4.0B≈2.6 GB at 4-bit2 also selling it hosted
- R-4BYannQi4.0B≈2.6 GB at 4-bit
- Voxtral-4B-TTS-2603Mistral AI4.0B≈2.6 GB at 4-bit
- Youtu-VL-4B-InstructTencent Hunyuan4.0B≈2.6 GB at 4-bit
- gemma-3-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- medgemma-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- shieldgemma-2-4b-itGoogle4.0B≈2.6 GB at 4-bit
- t5gemma-2-4b-4bGoogle4.0B≈2.6 GB at 4-bit
- OmniGen2OmniGen4.0B≈2.6 GB at 4-bit
- HeartMuLa-oss-3BHeartMuLa3.9B≈2.6 GB at 4-bit
- Nanbeige4.1-3BNanbeige LLM Lab3.9B≈2.6 GB at 4-bit
- Phi-4-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- LocateAnything-3BNVIDIA3.8B≈2.5 GB at 4-bit
- NuExtractNuMind3.8B≈2.5 GB at 4-bit
- NuExtract-1.5NuMind3.8B≈2.5 GB at 4-bit
- Phi-3-mini-128k-instructMicrosoft3.8B≈2.5 GB at 4-bit2 also selling it hosted
- Phi-3-mini-4k-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Phi-3.5-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Hy-Embodied-0.5Tencent Hunyuan3.8B≈2.5 GB at 4-bit
- Orpheus 3BCanopy Labs3.8B≈2.5 GB at 4-bit1 also selling it hosted
- blip2-opt-2.7bSalesforce3.7B≈2.4 GB at 4-bit
- HyperCLOVAX-SEED-Vision-Instruct-3BHyperCLOVA X3.7B≈2.4 GB at 4-bit
- Ovis-U1-3BATH-MaaS3.6B≈2.4 GB at 4-bit
- YuE2-3BMultimodal Art Projection3.6B≈2.4 GB at 4-bit
- Step-Audio-TTS-3Bstepfun-ai3.5B≈2.3 GB at 4-bit
- ACE-Step-v1-3.5BACE-Step3.5B≈2.3 GB at 4-bit
- NSFW-GEN-ANIMEUnfilteredAI3.5B≈2.3 GB at 4-bit
- NSFW-gen-v2UnfilteredAI3.5B≈2.3 GB at 4-bit
- Breeze-TTS-2BreezeBlue3.5B≈2.3 GB at 4-bit
- deepseek-vl2-tinydeepseek-ai3.4B≈2.2 GB at 4-bit
- Unlimited-OCRBaidu3.3B≈2.2 GB at 4-bit
- FLUX.1-dev-ControlNet-Union-ProShakker Labs3.3B≈2.1 GB at 4-bit
- nllb-200-3.3BAI at Meta3.3B≈2.1 GB at 4-bit
- Llama 3.2 3B InstructMeta3.2B≈2.1 GB at 4-bit14 also selling it hosted
- llama-3.2-Korean-Bllossom-3BBllossom3.2B≈2.1 GB at 4-bit
- Granite 4.0 MicroIBM watsonx.ai3.2B≈2.1 GB at 4-bit2 also selling it hosted
- imp-v1-3bMILVLG3.2B≈2.1 GB at 4-bit
- LFM2.5-VL-3BLiquid AI3.1B≈2.0 GB at 4-bit
- Qwen2.5-3BAlibaba3.1B≈2.0 GB at 4-bit
- Qwen2.5-3B-InstructAlibaba3.1B≈2.0 GB at 4-bit
- VibeThinker-3BWeiboAI3.1B≈2.0 GB at 4-bit
- SmolLM3-3BHugging Face Smol Models Research3.1B≈2.0 GB at 4-bit
- dots.ocrdots studio3.0B≈2.0 GB at 4-bit
- OpenELM-3B-InstructApple3.0B≈2.0 GB at 4-bit
- starcoder2-3bbigcode3.0B≈2.0 GB at 4-bit1 also selling it hosted
- Qwen2.5-Coder-3B-InstructAlibaba3.0B≈2.0 GB at 4-bit3 also selling it hosted
- SmolLM3-3B-BaseHuggingFaceTB3.0B≈2.0 GB at 4-bit