Models that run on 96 GB
1115 models with published weights that fit in 96 GB — a MacBook Pro with 96. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 102 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 1,115 in all, a hundred to a page; this is page 9 of 12.
An Apple configuration for people who intend to run models rather than occasionally try one. Comfortably holds the seventy-billion class with a very long context, or a mixture-of-experts model whose active parameters are few but whose weights are all resident.
- playground-v2.5-1024px-aestheticPlayground2.6B≈1.7 GB at 4-bit1 also selling it hosted
- sdxl-flashStable Diffusion Community (Unofficial, Non-profit)2.6B≈1.7 GB at 4-bit
- stable-diffusion-xl-base-0.9Stability AI2.6B≈1.7 GB at 4-bit
- MiniCPM5-2BOpenBMB2.5B≈1.6 GB at 4-bit
- Octopus-v2Nexa AI2.5B≈1.6 GB at 4-bit
- gemma-2bGoogle2.5B≈1.6 GB at 4-bit
- gemma-2b-itGoogle2.5B≈1.6 GB at 4-bit1 also selling it hosted
- BidirLM-Omni-2.5B-EmbeddingBidirLM2.5B≈1.6 GB at 4-bit
- canary-qwen-2.5bNVIDIA2.5B≈1.6 GB at 4-bit
- Cosmos-Reason2-2BNVIDIA2.4B≈1.6 GB at 4-bit
- EXAONE-3.5-2.4B-InstructLGAI-EXAONE2.4B≈1.6 GB at 4-bit
- seamless-m4t-v2-largeAI at Meta2.3B≈1.5 GB at 4-bit
- VoxCPM2OpenBMB2.3B≈1.5 GB at 4-bit
- stable-diffusion-xl-refiner-0.9Stability AI2.3B≈1.5 GB at 4-bit
- stable-diffusion-xl-refiner-1.0Stability AI2.3B≈1.5 GB at 4-bit
- GLM-ASR-Nano-2512Z.ai2.3B≈1.5 GB at 4-bit
- SmolVLM2-2.2B-InstructHugging Face Smol Models Research2.2B≈1.5 GB at 4-bit
- Marlin-2BNemo Station2.2B≈1.4 GB at 4-bit
- Qwen2-VL-2B-InstructAlibaba2.2B≈1.4 GB at 4-bit1 also selling it hosted
- tada-1bHume AI2.2B≈1.4 GB at 4-bit
- FLUX.1-dev-Controlnet-Inpainting-Betaalimama-creative2.1B≈1.4 GB at 4-bit
- FLUX.1-dev-ControlNet-Union-Pro-2.0Shakker Labs2.1B≈1.4 GB at 4-bit
- Qwen3-VL-2B-InstructAlibaba2.1B≈1.4 GB at 4-bit
- Qwen3-VL-Embedding-2BAlibaba2.1B≈1.4 GB at 4-bit
- Janus-1.3BDeepSeek2.1B≈1.4 GB at 4-bit
- stable-diffusion-3-medium-diffusersStability AI2.1B≈1.4 GB at 4-bit
- cohere-transcribe-arabic-07-2026Cohere Labs2.1B≈1.3 GB at 4-bit
- SoulX-Podcast-1.7BSoul-AILab2.1B≈1.3 GB at 4-bit
- Qwen3-1.7BAlibaba2.0B≈1.3 GB at 4-bit1 also selling it hosted
- Isaac 0.2 2B PreviewPerceptron2.0B≈1.3 GB at 4-bit1 also selling it hosted
- MOSS-Transcribe-preview-2BOpenMOSS-Team2.0B≈1.3 GB at 4-bit
- MiniCPM-2B-sft-fp32OpenBMB2.0B≈1.3 GB at 4-bit
- Qwen3.5-2BAlibaba2.0B≈1.3 GB at 4-bit2 also selling it hosted
- gemma-1.1-2b-itGoogle2.0B≈1.3 GB at 4-bit
- granite-speech-4.1-2bIBM watsonx.ai2.0B≈1.3 GB at 4-bit
- granite-speech-4.1-2b-narIBM watsonx.ai2.0B≈1.3 GB at 4-bit
- helium-1-preview-2bKyutai2.0B≈1.3 GB at 4-bit
- Youtu-LLM-2BTencent Hunyuan2.0B≈1.3 GB at 4-bit
- moondream2vikhyatk1.9B≈1.3 GB at 4-bit
- LTX-VideoLTX.io1.9B≈1.3 GB at 4-bit
- Dia2-2BNari Labs1.9B≈1.2 GB at 4-bit
- Qwen3-TTS-12Hz-1.7B-CustomVoiceAlibaba1.9B≈1.2 GB at 4-bit
- Qwen3-TTS-12Hz-1.7B-VoiceDesignAlibaba1.9B≈1.2 GB at 4-bit
- Hy-MT2-1.8BTencent Hunyuan1.8B≈1.2 GB at 4-bit1 also selling it hosted
- Flux.1-dev-Controlnet-UpscalerJasper.ai1.8B≈1.2 GB at 4-bit
- DeepScaleR-1.5B-PreviewAgentica1.8B≈1.2 GB at 4-bit
- DeepSeek-R1-Distill-Qwen-1.5BDeepSeek1.8B≈1.2 GB at 4-bit3 also selling it hosted
- Nemotron-Research-Reasoning-Qwen-1.5BNVIDIA1.8B≈1.2 GB at 4-bit
- VibeThinker-1.5BWeiboAI1.8B≈1.2 GB at 4-bit
- gte-Qwen2-1.5B-instructAlibaba-NLP1.8B≈1.2 GB at 4-bit
- Qwen3.6-27B-DFlashZ Lab1.7B≈1.1 GB at 4-bit
- BidirLM-1.7B-EmbeddingBidirLM1.7B≈1.1 GB at 4-bit
- F2LLM-1.7Bcodefuse-ai1.7B≈1.1 GB at 4-bit
- F2LLM-v2-1.7Bcodefuse-ai1.7B≈1.1 GB at 4-bit
- Qwen3-ASR-1.7BAlibaba1.7B≈1.1 GB at 4-bit1 also selling it hosted
- SmolLM-1.7BHuggingFaceTB1.7B≈1.1 GB at 4-bit
- SmolLM2-1.7BHuggingFaceTB1.7B≈1.1 GB at 4-bit
- CogVideoX-2bZ.ai1.7B≈1.1 GB at 4-bit
- stablelm-2-1_6bStability AI1.6B≈1.1 GB at 4-bit
- stablelm-2-zephyr-1_6bStability AI1.6B≈1.1 GB at 4-bit
- Dia-1.6BNari Labs1.6B≈1.0 GB at 4-bit
- CrisperWhispernyra labs1.6B≈1.0 GB at 4-bit
- gpt2-xlOpenAI community1.6B≈1.0 GB at 4-bit
- LFM2.5-VL-1.6BLiquid AI1.6B≈1.0 GB at 4-bit
- tts-1.6b-en_frKyutai1.6B≈1.0 GB at 4-bit
- LFM2-VL-1.6BLiquid AI1.6B≈1.0 GB at 4-bit
- musicgen-melodyfacebook1.6B≈1.0 GB at 4-bit
- Qwen2.5-1.5BAlibaba1.5B≈1.0 GB at 4-bit
- Qwen2.5-1.5B-InstructAlibaba1.5B≈1.0 GB at 4-bit1 also selling it hosted
- ReaderLM-v2Jina AI1.5B≈1.0 GB at 4-bit
- reader-lm-1.5bJina AI1.5B≈1.0 GB at 4-bit
- Whisper Large v3OpenAI1.5B≈1.0 GB at 4-bit2 also selling it hosted
- whisper-largeOpenAI1.5B≈1.0 GB at 4-bit
- whisper-large-v2OpenAI1.5B≈1.0 GB at 4-bit
- stable-video-diffusion-img2vidStability AI1.5B≈1.0 GB at 4-bit
- stable-video-diffusion-img2vid-xtStability AI1.5B≈1.0 GB at 4-bit
- stable-video-diffusion-img2vid-xt-1-1Stability AI1.5B≈1.0 GB at 4-bit
- Hymba-1.5B-InstructNVIDIA1.5B≈1.0 GB at 4-bit
- Arch-Router-1.5Bkatanemo1.5B≈1.0 GB at 4-bit
- Hymba-1.5B-BaseNVIDIA1.5B≈1.0 GB at 4-bit
- HyperCLOVAX-SEED-Text-Instruct-1.5BHyperCLOVA X1.5B≈1.0 GB at 4-bit
- Qwen2-1.5B-InstructAlibaba1.5B≈1.0 GB at 4-bit1 also selling it hosted
- stella_en_1.5B_v5NovaSearch1.5B≈1.0 GB at 4-bit
- LFM2-Audio-1.5BLiquid AI1.5B≈1.0 GB at 4-bit
- LFM2.5-Audio-1.5BLiquid AI1.5B≈1.0 GB at 4-bit
- starvector-1b-im2svgstarvector1.4B≈0.9 GB at 4-bit
- i2vgen-xlali-vilab1.4B≈0.9 GB at 4-bit
- phi-1Microsoft1.4B≈0.9 GB at 4-bit
- phi-1_5Microsoft1.4B≈0.9 GB at 4-bit
- text-to-video-ms-1.7bali-vilab1.4B≈0.9 GB at 4-bit
- Ouro-1.4BByteDance1.4B≈0.9 GB at 4-bit
- SSD-1BSegmind1.3B≈0.9 GB at 4-bit1 also selling it hosted
- MiniCPM-V-4.6OpenBMB1.3B≈0.8 GB at 4-bit
- FastWan-QAD-FP8-1.3BFastVideo1.3B≈0.8 GB at 4-bit1 also selling it hosted
- JanusFlow-1.3Bdeepseek-ai1.3B≈0.8 GB at 4-bit
- Wan2.1-T2V-1.3BWan-AI1.3B≈0.8 GB at 4-bit1 also selling it hosted
- Wan2.1-T2V-1.3B-DiffusersWan-AI1.3B≈0.8 GB at 4-bit
- Wan2.1-VACE-1.3BWan-AI1.3B≈0.8 GB at 4-bit
- deepseek-coder-1.3b-instructdeepseek-ai1.3B≈0.8 GB at 4-bit
- gpt-neo-1.3BEleutherAI1.3B≈0.8 GB at 4-bit