Models that run on 24 GB
890 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 890 in all, a hundred to a page; this is page 5 of 9.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
- stablelm-base-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- stablelm-tuned-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- vicuna-7b-v1.5Large Model Systems Organization7.0B≈4.5 GB at 4-bit
- xgen-7b-8k-baseSalesforce7.0B≈4.5 GB at 4-bit
- granite-4.0-h-tinyIBM Granite6.9B≈4.5 GB at 4-bit
- NuminaMath-7B-TIRProject-Numina6.9B≈4.5 GB at 4-bit
- OLMo-7BAi26.9B≈4.5 GB at 4-bit
- Magicoder-S-DS-6.7BIntellligent Software Engineering (iSE)6.7B≈4.4 GB at 4-bit
- deepseek-coder-6.7b-instructDeepSeek6.7B≈4.4 GB at 4-bit
- meditron-7bEPFL LLM Team6.7B≈4.4 GB at 4-bit
- CodeLlama-7b-Instruct-hfCode Llama6.7B≈4.4 GB at 4-bit
- CodeLlama-7b-hfCode Llama6.7B≈4.4 GB at 4-bit
- sqlcoder-7b-2Defog.ai6.7B≈4.4 GB at 4-bit
- Llama-2-7b-chat-hfMeta Llama6.7B≈4.4 GB at 4-bit
- Llama-2-7b-hfMeta Llama6.7B≈4.4 GB at 4-bit
- llama-7bhuggyllama6.7B≈4.4 GB at 4-bit
- llama2_7b_chat_uncensoredgeorgesung6.7B≈4.4 GB at 4-bit
- LlamaGuard-7bMeta Llama6.7B≈4.4 GB at 4-bit1 also selling it hosted
- dinov3-vit7b16-pretrain-lvd1689mfacebook6.7B≈4.4 GB at 4-bit
- YuE-s1-7B-anneal-en-cotMultimodal Art Projection6.2B≈4.0 GB at 4-bit
- Z-ImageTongyi-MAI6.2B≈4.0 GB at 4-bit
- Z-Image-TurboTongyi-MAI6.2B≈4.0 GB at 4-bit
- Yi-6B01-ai6.1B≈3.9 GB at 4-bit1 also selling it hosted
- Yi-6B-200K01-ai6.0B≈3.9 GB at 4-bit
- gpt-j-6bEleutherAI6.0B≈3.9 GB at 4-bit
- pygmalion-6bPygmalion6.0B≈3.9 GB at 4-bit
- Chroma-4BFlashLabs5.9B≈3.8 GB at 4-bit
- higgs-tts-2-3b-baseBoson AI5.8B≈3.8 GB at 4-bit
- DeciLM-6bDeci AI5.7B≈3.7 GB at 4-bit
- CogVideoX-5bZ.ai5.6B≈3.6 GB at 4-bit
- Qwen2.5-Omni-3BAlibaba5.5B≈3.6 GB at 4-bit
- chandra-ocr-2Datalab5.3B≈3.4 GB at 4-bit
- Gemma-4-E2B-itGoogle5.1B≈3.3 GB at 4-bit1 also selling it hosted
- CogVideoX-5b-I2Vzai-org5.0B≈3.2 GB at 4-bit
- CogVideoX1.5-5B-SATzai-org5.0B≈3.2 GB at 4-bit
- Wan2.2-TI2V-5BWan-AI5.0B≈3.2 GB at 4-bit
- Wan2.2-TI2V-5B-DiffusersWan-AI5.0B≈3.2 GB at 4-bit
- acestep-v15-xl-turboACE-Step5.0B≈3.2 GB at 4-bit
- translategemma-4b-itGoogle5.0B≈3.2 GB at 4-bit
- Mage-VLMicrosoft4.7B≈3.1 GB at 4-bit
- higgs-tts-3-4bBoson AI4.7B≈3.0 GB at 4-bit
- s2-proFish Audio4.6B≈3.0 GB at 4-bit1 also selling it hosted
- NuExtract3NuMind4.5B≈3.0 GB at 4-bit
- Qwen-Drive-1.0-4BAlibaba4.5B≈3.0 GB at 4-bit
- Llama-3.1-Minitron-4B-Width-BaseNVIDIA4.5B≈2.9 GB at 4-bit
- Voxtral-Mini-4B-Realtime-2602Mistral AI4.4B≈2.9 GB at 4-bit
- IF-I-XL-v1.0DeepFloyd4.3B≈2.8 GB at 4-bit
- Gemma 3 4BGoogle4.3B≈2.8 GB at 4-bit6 also selling it hosted
- medgemma-1.5-4b-itGoogle4.3B≈2.8 GB at 4-bit
- medgemma-4b-itGoogle4.3B≈2.8 GB at 4-bit
- NeoHorse-1-4BTokenRhythm4.2B≈2.7 GB at 4-bit
- Nanbeige4.2-3BNanbeige LLM Lab4.2B≈2.7 GB at 4-bit
- Phi-3-vision-128k-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Phi-3.5-vision-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Spark-X2.5-4BSparkLLM4.1B≈2.7 GB at 4-bit
- MiniCPM-V-4OpenBMB4.1B≈2.6 GB at 4-bit
- AgentCPM-ExploreOpenBMB4.0B≈2.6 GB at 4-bit
- DASD-4B-ThinkingAlibaba Cloud Apsara Lab4.0B≈2.6 GB at 4-bit
- Jan-nanoMenlo Research4.0B≈2.6 GB at 4-bit
- Jan-nano-128kMenlo Research4.0B≈2.6 GB at 4-bit
- Jan-v1-4BJan4.0B≈2.6 GB at 4-bit
- fable-tracesAliesTaha4.0B≈2.6 GB at 4-bit
- Mellum-4b-baseJetBrains4.0B≈2.6 GB at 4-bit
- Llasa-3BHKUST Audio4.0B≈2.6 GB at 4-bit
- F2LLM-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- F2LLM-v2-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- FLUX.2-klein-base-4BBlack Forest Labs4.0B≈2.6 GB at 4-bit
- GELab-Zero-4B-previewstepfun-ai4.0B≈2.6 GB at 4-bit
- Gemma-3-Gaia-PT-BR-4b-itCEIA-UFG4.0B≈2.6 GB at 4-bit
- LocoOperator-4BLocoreMind4.0B≈2.6 GB at 4-bit
- LocoTrainer-4BLocoreMind4.0B≈2.6 GB at 4-bit
- MiniCPM3-4BOpenBMB4.0B≈2.6 GB at 4-bit
- Nemotron-Mini-4B-InstructNVIDIA4.0B≈2.6 GB at 4-bit
- OmniNeural-4BNexa AI4.0B≈2.6 GB at 4-bit
- Qwen3 4BAlibaba4.0B≈2.6 GB at 4-bit3 also selling it hosted
- Qwen3-4B-Instruct-2507Alibaba4.0B≈2.6 GB at 4-bit2 also selling it hosted
- Qwen3-4B-Thinking-2507Alibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3-4b-Z-Image-Engineer-V4BennyDaBall4.0B≈2.6 GB at 4-bit
- Qwen3-Embedding-4BAlibaba4.0B≈2.6 GB at 4-bit4 also selling it hosted
- Qwen3-Reranker-4BAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3-VL-4B-InstructAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3-VL-4B-ThinkingAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3.5-4BAlibaba4.0B≈2.6 GB at 4-bit2 also selling it hosted
- R-4BYannQi4.0B≈2.6 GB at 4-bit
- Voxtral-4B-TTS-2603Mistral AI4.0B≈2.6 GB at 4-bit
- Youtu-VL-4B-InstructTencent Hunyuan4.0B≈2.6 GB at 4-bit
- gemma-3-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- medgemma-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- shieldgemma-2-4b-itGoogle4.0B≈2.6 GB at 4-bit
- t5gemma-2-4b-4bGoogle4.0B≈2.6 GB at 4-bit
- OmniGen2OmniGen4.0B≈2.6 GB at 4-bit
- HeartMuLa-oss-3BHeartMuLa3.9B≈2.6 GB at 4-bit
- Nanbeige4.1-3BNanbeige LLM Lab3.9B≈2.6 GB at 4-bit
- Phi-4-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- LocateAnything-3BNVIDIA3.8B≈2.5 GB at 4-bit
- NuExtractNuMind3.8B≈2.5 GB at 4-bit
- NuExtract-1.5NuMind3.8B≈2.5 GB at 4-bit
- Phi-3-mini-128k-instructMicrosoft3.8B≈2.5 GB at 4-bit2 also selling it hosted
- Phi-3-mini-4k-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Phi-3.5-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted