Conversation models that run on 24 GB
455 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 455 in all, a hundred to a page; this is page 1 of 5.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- Dolphin Mistral 24B Venice EditionCognitive Computations24.0B≈15.6 GB at 4-bit1 also selling it hosted
- Mistral Small 3.1 24B Instruct 2503Mistral AI24.0B≈15.6 GB at 4-bit3 also selling it hosted
- Mistral-Small-3.2-24B-Instruct-2506Mistral AI24.0B≈15.6 GB at 4-bit1 also selling it hosted
- Mistral Small 3Mistral AI23.6B≈15.3 GB at 4-bit4 also selling it hosted
- sarvam-mSarvam AI23.6B≈15.3 GB at 4-bit
- solar-pro-preview-instructUpstage22.1B≈14.4 GB at 4-bit
- gpt-oss-safeguard-20bOpenAI21.5B≈14.0 GB at 4-bit6 also selling it hosted
- ERNIE-4.5-21B-A3B-PTBaidu21.0B≈13.7 GB at 4-bit1 also selling it hosted
- context-1chroma20.9B≈13.6 GB at 4-bit
- gpt-neox-20bEleutherAI20.7B≈13.5 GB at 4-bit
- maple-previewdeepgrove20.2B≈13.1 GB at 4-bit
- Gemma-4-31B-JANG_4M-CRACKdealignai20.2B≈13.1 GB at 4-bit
- cogvlm2-llama3-chat-19Bzai-org19.5B≈12.7 GB at 4-bit
- cogvlm-chat-hfzai-org17.6B≈11.5 GB at 4-bit
- Ling-mini-2.0InclusionAI16.3B≈10.6 GB at 4-bit
- Ring-mini-2.0InclusionAI16.3B≈10.6 GB at 4-bit
- Instella-MoE-16B-A3B-Thinkamd16.0B≈10.4 GB at 4-bit
- deepseek-moe-16b-basedeepseek-ai16.0B≈10.4 GB at 4-bit
- deepseek-moe-16b-chatdeepseek-ai16.0B≈10.4 GB at 4-bit
- Moonlight-16B-A3B-InstructMoonshot AI16.0B≈10.4 GB at 4-bit
- DeepSeek-V2-Litedeepseek-ai15.7B≈10.2 GB at 4-bit
- starchat-alphaHugging Face H415.5B≈10.1 GB at 4-bit
- Apriel-1.6-15b-ThinkerServiceNow-AI15.0B≈9.8 GB at 4-bit
- Apriel-1.5-15b-ThinkerServiceNow-AI14.9B≈9.7 GB at 4-bit
- Qwen2.5-14B-Instruct-1MAlibaba14.8B≈9.6 GB at 4-bit
- SuperNova-MediusArcee AI14.8B≈9.6 GB at 4-bit
- Qwen1.5-MoE-A2.7BAlibaba14.3B≈9.3 GB at 4-bit
- 14BCausalLM14.0B≈9.1 GB at 4-bit
- ChatTS-14Bbytedance-research14.0B≈9.1 GB at 4-bit
- Fathom-R1-14BFractalAIResearch14.0B≈9.1 GB at 4-bit
- Nemotron-Labs-Diffusion-14BNVIDIA14.0B≈9.1 GB at 4-bit
- Qwen2.5-14BAlibaba14.0B≈9.1 GB at 4-bit1 also selling it hosted
- Velvet-14BAlmawave14.0B≈9.1 GB at 4-bit
- rwkv-4-pile-14bBlinkDL14.0B≈9.1 GB at 4-bit
- miniGCausalLM14.0B≈9.1 GB at 4-bit
- NexusRaven-V2-13BNexusflow13.0B≈8.5 GB at 4-bit
- Llama-2-13b-hfMeta Llama13.0B≈8.5 GB at 4-bit
- Baichuan2-13B-ChatBaichuan Intelligent Technology13.0B≈8.5 GB at 4-bit
- LLaVA-13b-delta-v0liuhaotian13.0B≈8.5 GB at 4-bit
- Llama-2-13bMeta Llama13.0B≈8.5 GB at 4-bit
- Llama-2-13b-chatMeta Llama13.0B≈8.5 GB at 4-bit1 also selling it hosted
- Llama-2-13b-chat-hfMeta13.0B≈8.5 GB at 4-bit1 also selling it hosted
- Llama2-13B-TiefighterKoboldAI13.0B≈8.5 GB at 4-bit1 also selling it hosted
- Llama2-Chinese-13b-ChatFlagAlpha13.0B≈8.5 GB at 4-bit
- MythoMax 13BGryphe13.0B≈8.5 GB at 4-bit4 also selling it hosted
- Nous Hermes Llama2 13B13.0B≈8.5 GB at 4-bit2 also selling it hosted
- OpenOrca-Platypus2-13BOpenOrca13.0B≈8.5 GB at 4-bit
- Orca-2-13bMicrosoft13.0B≈8.5 GB at 4-bit
- ReMM SLERP 13BUndi9513.0B≈8.5 GB at 4-bit1 also selling it hosted
- WhiteRabbitNeo-13B-v1WhiteRabbitNeo13.0B≈8.5 GB at 4-bit
- Wizard-Vicuna-13B-UncensoredQuixi AI13.0B≈8.5 GB at 4-bit
- Wizard-Vicuna-13B-Uncensored-HFTheBloke13.0B≈8.5 GB at 4-bit
- WizardLM-13B-UncensoredQuixi AI13.0B≈8.5 GB at 4-bit
- WizardLM-13B-V1.2WizardLM Team13.0B≈8.5 GB at 4-bit
- Ziya-LLaMA-13B-v1Fengshenbang-LM13.0B≈8.5 GB at 4-bit
- chronos-hermes-13b-v2Austism13.0B≈8.5 GB at 4-bit2 also selling it hosted
- jais-13binception4213.0B≈8.5 GB at 4-bit
- jais-13b-chatinception4213.0B≈8.5 GB at 4-bit1 also selling it hosted
- llama-13bhuggyllama13.0B≈8.5 GB at 4-bit
- mythalion-13bPygmalionAI13.0B≈8.5 GB at 4-bit
- open_llama_13bOpenLM Research13.0B≈8.5 GB at 4-bit
- prometheus-13b-v1.0prometheus-eval13.0B≈8.5 GB at 4-bit
- ruGPT-3.5-13Bai-forever13.0B≈8.5 GB at 4-bit
- stable-vicuna-13b-deltaCarperAI13.0B≈8.5 GB at 4-bit
- vicuna-13b-v1.5Large Model Systems Organization13.0B≈8.5 GB at 4-bit
- vicuna-13b-v1.5-16kLarge Model Systems Organization13.0B≈8.5 GB at 4-bit
- Wayfarer-12BLatitude12.2B≈8.0 GB at 4-bit
- Mistral NemoMistral AI12.2B≈8.0 GB at 4-bit10 also selling it hosted
- MN-12B-Celeste-V1.9nothingiisreal12.0B≈7.8 GB at 4-bit
- NVIDIA-Nemotron-Nano-12B-v2NVIDIA12.0B≈7.8 GB at 4-bit2 also selling it hosted
- NemoMix-Unleashed-12BMarinaraSpaghetti12.0B≈7.8 GB at 4-bit
- Pixtral 12B 2409Mistral AI12.0B≈7.8 GB at 4-bit3 also selling it hosted
- gemma-3-12b-it-qat-q4_0-unquantizedGoogle12.0B≈7.8 GB at 4-bit
- oasst-sft-1-pythia-12bOpenAssistant12.0B≈7.8 GB at 4-bit
- oasst-sft-4-pythia-12b-epoch-3.5OpenAssistant12.0B≈7.8 GB at 4-bit
- pythia-12bEleutherAI12.0B≈7.8 GB at 4-bit1 also selling it hosted
- falcon-11BTechnology Innovation Institute11.1B≈7.2 GB at 4-bit
- Bielik-11B-v3.0-Instructspeakleash11.0B≈7.2 GB at 4-bit1 also selling it hosted
- Llama 3.2 11B Vision InstructMeta11.0B≈7.2 GB at 4-bit7 also selling it hosted
- Llama-3.2V-11B-cotXkev11.0B≈7.2 GB at 4-bit
- YanoljaNEXT-EEVE-Instruct-10.8Byanolja10.8B≈7.0 GB at 4-bit
- HyperCLOVAX-SEED-Omni-8BHyperCLOVA X10.7B≈7.0 GB at 4-bit
- Fimbulvetr-11B-v2Sao10K10.7B≈7.0 GB at 4-bit
- Solar-10.7B-Instruct-v1.0Upstage10.7B≈7.0 GB at 4-bit
- Solar-10.7B-v1.0Upstage10.7B≈7.0 GB at 4-bit
- Llama-3.2-11B-VisionMeta Llama10.6B≈6.9 GB at 4-bit
- Step3-VL-10BStepFun10.2B≈6.6 GB at 4-bit
- Qwen3.8-9B-Distillempero-ai9.7B≈6.3 GB at 4-bit
- fuyu-8bAdept AI Labs9.4B≈6.1 GB at 4-bit
- moondream3-previewmoondream9.3B≈6.0 GB at 4-bit
- gemma-2-9bGoogle9.2B≈6.0 GB at 4-bit
- EuroLLM-9B-InstructUTTER - Unified Transcription and Translation for Extended Reality9.2B≈5.9 GB at 4-bit
- EuroLLM-9BUTTER - Unified Transcription and Translation for Extended Reality9.0B≈5.9 GB at 4-bit
- Ornith-1.5-9Bornith-ai9.0B≈5.9 GB at 4-bit
- Ovis1.6-Gemma2-9BATH-MaaS9.0B≈5.9 GB at 4-bit
- Ovis2.5-9BATH-MaaS9.0B≈5.9 GB at 4-bit
- Qwen3.5-9B-BaseAlibaba9.0B≈5.9 GB at 4-bit1 also selling it hosted
- Qwythos-9B-v2empero-ai9.0B≈5.9 GB at 4-bit
- Turkish-Gemma-9b-T1ytu-ce-cosmos9.0B≈5.9 GB at 4-bit
- Yi-1.5-9B-Chat01-ai9.0B≈5.9 GB at 4-bit