Conversation models that run on 24 GB
455 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 455 in all, a hundred to a page; this is page 4 of 5.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- Llama 3.2 3B InstructMeta3.2B≈2.1 GB at 4-bit14 also selling it hosted
- llama-3.2-Korean-Bllossom-3BBllossom3.2B≈2.1 GB at 4-bit
- imp-v1-3bMILVLG3.2B≈2.1 GB at 4-bit
- LFM2.5-VL-3BLiquid AI3.1B≈2.0 GB at 4-bit
- Qwen2.5-3BAlibaba3.1B≈2.0 GB at 4-bit
- Qwen2.5-3B-InstructAlibaba3.1B≈2.0 GB at 4-bit
- VibeThinker-3BWeiboAI3.1B≈2.0 GB at 4-bit
- SmolLM3-3BHugging Face Smol Models Research3.1B≈2.0 GB at 4-bit
- OpenELM-3B-InstructApple3.0B≈2.0 GB at 4-bit
- SmolLM3-3B-BaseHuggingFaceTB3.0B≈2.0 GB at 4-bit
- bitnet_b1_58-3B1bitLLM3.0B≈2.0 GB at 4-bit
- open_llama_3bOpenLM Research3.0B≈2.0 GB at 4-bit
- open_llama_3b_v2OpenLM Research3.0B≈2.0 GB at 4-bit
- orca_mini_3bpankajmathur3.0B≈2.0 GB at 4-bit
- paligemma2-3b-pt-224Google3.0B≈2.0 GB at 4-bit
- proxy-lite-3bconvergence-ai3.0B≈2.0 GB at 4-bit
- stablelm-3b-4e1tStability AI3.0B≈2.0 GB at 4-bit
- paligemma-3b-pt-224Google2.9B≈1.9 GB at 4-bit
- stablelm-zephyr-3bStability AI2.8B≈1.8 GB at 4-bit
- dolphin-2_6-phi-2Dolphin2.8B≈1.8 GB at 4-bit
- phi-2Microsoft2.8B≈1.8 GB at 4-bit
- Solidity-LLMChainGPT2.8B≈1.8 GB at 4-bit
- gpt-neo-2.7BEleutherAI2.7B≈1.8 GB at 4-bit
- gemma-2-2bGoogle2.6B≈1.7 GB at 4-bit
- gemma-2-2b-itGoogle2.6B≈1.7 GB at 4-bit
- gemma-2-2b-jpn-itGoogle2.6B≈1.7 GB at 4-bit
- LFM2-2.6B-TranscriptLiquid AI2.6B≈1.7 GB at 4-bit
- Ovis2.5-2BATH-MaaS2.6B≈1.7 GB at 4-bit
- LFM2-2.6BLiquid AI2.6B≈1.7 GB at 4-bit
- LFM2-2.6B-ExpLiquid AI2.6B≈1.7 GB at 4-bit
- MiniCPM5-2BOpenBMB2.5B≈1.6 GB at 4-bit
- Octopus-v2Nexa AI2.5B≈1.6 GB at 4-bit
- gemma-2bGoogle2.5B≈1.6 GB at 4-bit
- gemma-2b-itGoogle2.5B≈1.6 GB at 4-bit1 also selling it hosted
- Cosmos-Reason2-2BNVIDIA2.4B≈1.6 GB at 4-bit
- EXAONE-3.5-2.4B-InstructLGAI-EXAONE2.4B≈1.6 GB at 4-bit
- SmolVLM2-2.2B-InstructHugging Face Smol Models Research2.2B≈1.5 GB at 4-bit
- Marlin-2BNemo Station2.2B≈1.4 GB at 4-bit
- Qwen2-VL-2B-InstructAlibaba2.2B≈1.4 GB at 4-bit1 also selling it hosted
- Qwen3-VL-2B-InstructAlibaba2.1B≈1.4 GB at 4-bit
- MiniCPM-2B-sft-fp32OpenBMB2.0B≈1.3 GB at 4-bit
- Qwen3.5-2BAlibaba2.0B≈1.3 GB at 4-bit2 also selling it hosted
- gemma-1.1-2b-itGoogle2.0B≈1.3 GB at 4-bit
- helium-1-preview-2bKyutai2.0B≈1.3 GB at 4-bit
- Youtu-LLM-2BTencent Hunyuan2.0B≈1.3 GB at 4-bit
- moondream2vikhyatk1.9B≈1.3 GB at 4-bit
- VibeThinker-1.5BWeiboAI1.8B≈1.2 GB at 4-bit
- Qwen3.6-27B-DFlashZ Lab1.7B≈1.1 GB at 4-bit
- SmolLM-1.7BHuggingFaceTB1.7B≈1.1 GB at 4-bit
- SmolLM2-1.7BHuggingFaceTB1.7B≈1.1 GB at 4-bit
- stablelm-2-1_6bStability AI1.6B≈1.1 GB at 4-bit
- stablelm-2-zephyr-1_6bStability AI1.6B≈1.1 GB at 4-bit
- LFM2.5-VL-1.6BLiquid AI1.6B≈1.0 GB at 4-bit
- LFM2-VL-1.6BLiquid AI1.6B≈1.0 GB at 4-bit
- Qwen2.5-1.5BAlibaba1.5B≈1.0 GB at 4-bit
- Qwen2.5-1.5B-InstructAlibaba1.5B≈1.0 GB at 4-bit1 also selling it hosted
- ReaderLM-v2Jina AI1.5B≈1.0 GB at 4-bit
- reader-lm-1.5bJina AI1.5B≈1.0 GB at 4-bit
- Hymba-1.5B-InstructNVIDIA1.5B≈1.0 GB at 4-bit
- Arch-Router-1.5Bkatanemo1.5B≈1.0 GB at 4-bit
- Hymba-1.5B-BaseNVIDIA1.5B≈1.0 GB at 4-bit
- HyperCLOVAX-SEED-Text-Instruct-1.5BHyperCLOVA X1.5B≈1.0 GB at 4-bit
- Qwen2-1.5B-InstructAlibaba1.5B≈1.0 GB at 4-bit1 also selling it hosted
- starvector-1b-im2svgstarvector1.4B≈0.9 GB at 4-bit
- phi-1Microsoft1.4B≈0.9 GB at 4-bit
- phi-1_5Microsoft1.4B≈0.9 GB at 4-bit
- Ouro-1.4BByteDance1.4B≈0.9 GB at 4-bit
- MiniCPM-V-4.6OpenBMB1.3B≈0.8 GB at 4-bit
- gpt-neo-1.3BEleutherAI1.3B≈0.8 GB at 4-bit
- nllb-200-distilled-1.3Bfacebook1.3B≈0.8 GB at 4-bit
- opt-1.3bfacebook1.3B≈0.8 GB at 4-bit
- EXAONE-4.0-1.2BLGAI-EXAONE1.3B≈0.8 GB at 4-bit
- SpatialLM-Llama-1BManycore Research1.2B≈0.8 GB at 4-bit
- LFM2.5-1.2B-BaseLiquid AI1.2B≈0.8 GB at 4-bit
- LFM2.5-1.2B-JPLiquid AI1.2B≈0.8 GB at 4-bit
- HRM-Text-1BSapient AI1.2B≈0.8 GB at 4-bit
- LFM2-1.2BLiquid AI1.2B≈0.8 GB at 4-bit
- LFM2.5-1.2B-InstructLiquid AI1.2B≈0.8 GB at 4-bit
- MinerU2.5-2509-1.2BOpenDataLab1.2B≈0.8 GB at 4-bit
- TinyLlama-1.1B-Chat-v1.0TinyLlama1.1B≈0.7 GB at 4-bit
- TinyLlama-1.1B-intermediate-step-1431k-3TTinyLlama1.1B≈0.7 GB at 4-bit
- MiniCPM5-1BOpenBMB1.1B≈0.7 GB at 4-bit
- vaultgemma-1bGoogle1.0B≈0.7 GB at 4-bit
- Gemma3-1B-ITLiteRT Community (FKA TFLite)1.0B≈0.7 GB at 4-bit
- MolmoE-1B-0924Allen Institute for AI (Ai2)1.0B≈0.7 GB at 4-bit
- antares-1bfdtn-ai1.0B≈0.7 GB at 4-bit
- granite-4.0-h-1bIBM Granite1.0B≈0.7 GB at 4-bit
- gemma-3-1b-itGoogle1.0B≈0.6 GB at 4-bit
- gemma-3-1b-ptGoogle1.0B≈0.6 GB at 4-bit
- CLIP-ViT-H-14-laion2B-s32B-b79KLAION eV1.0B≈0.6 GB at 4-bit
- MobileLLM-R1-950MAI at Meta0.9B≈0.6 GB at 4-bit
- siglip-so400m-patch14-384Google0.9B≈0.6 GB at 4-bit
- bitnet-b1.58-2B-4TMicrosoft0.8B≈0.6 GB at 4-bit
- gpt2-largeOpenAI community0.8B≈0.5 GB at 4-bit
- t5gemma-2-270m-270mGoogle0.8B≈0.5 GB at 4-bit
- Florence-2-largeMicrosoft0.8B≈0.5 GB at 4-bit
- Florence-2-large-ftMicrosoft0.8B≈0.5 GB at 4-bit
- FastVLM-0.5BApple0.8B≈0.5 GB at 4-bit
- Qwen3-0.6BAlibaba0.8B≈0.5 GB at 4-bit1 also selling it hosted
- Qwen3-0.6B-BaseAlibaba0.6B≈0.4 GB at 4-bit