Conversation models that run on 24 GB
455 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 455 in all, a hundred to a page; this is page 5 of 5.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- bloom-560mBigScience Workshop0.6B≈0.4 GB at 4-bit
- Qwen1.5-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- reader-lm-0.5bJina AI0.5B≈0.3 GB at 4-bit
- Qwen2-0.5B-InstructAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2.5-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2.5-0.5B-InstructAlibaba0.5B≈0.3 GB at 4-bit1 also selling it hosted
- blip-image-captioning-largeSalesforce0.5B≈0.3 GB at 4-bit
- LFM2.5-VL-450MLiquid AI0.4B≈0.3 GB at 4-bit
- PairRMLLM Blender0.4B≈0.3 GB at 4-bit
- CLIP-GmP-ViT-L-14zer0int0.4B≈0.3 GB at 4-bit
- circuit-sparsityOpenAI0.4B≈0.3 GB at 4-bit
- bart-large-cnnAI at Meta0.4B≈0.3 GB at 4-bit
- ModernBERT-largeAnswer.AI0.4B≈0.3 GB at 4-bit
- blip-vqa-baseSalesforce0.4B≈0.3 GB at 4-bit
- LFM2-350MLiquid AI0.4B≈0.2 GB at 4-bit
- LFM2.5-350MLiquid AI0.4B≈0.2 GB at 4-bit
- nougat-basefacebook0.3B≈0.2 GB at 4-bit
- bert-large-uncased-whole-word-masking-finetuned-squadBERT community0.3B≈0.2 GB at 4-bit
- mdeberta-v3-base-squad2timpal0l0.3B≈0.2 GB at 4-bit
- functiongemma-270m-itGoogle0.3B≈0.2 GB at 4-bit
- gemma-3-270mGoogle0.3B≈0.2 GB at 4-bit
- gemma-3-270m-itGoogle0.3B≈0.2 GB at 4-bit
- granite-docling-258MIBM Granite0.3B≈0.2 GB at 4-bit
- Florence-2-baseMicrosoft0.2B≈0.2 GB at 4-bit
- t5-baseT5 community0.2B≈0.1 GB at 4-bit
- chatgpt_paraphraser_on_T5_baseHumarin0.2B≈0.1 GB at 4-bit
- BiRefNetZhengPeng70.2B≈0.1 GB at 4-bit
- RMBG-2.0BRIA AI0.2B≈0.1 GB at 4-bit
- bert-base-multilingual-casedBERT community0.2B≈0.1 GB at 4-bit
- gpt-neo-125mEleutherAI0.2B≈0.1 GB at 4-bit
- ModernBERT-baseAnswer.AI0.1B≈0.1 GB at 4-bit
- MagicPrompt-Stable-DiffusionGustavosta0.1B≈0.1 GB at 4-bit
- gpt2OpenAI community0.1B≈0.1 GB at 4-bit
- distilbert-base-multilingual-caseddistilbert0.1B≈0.1 GB at 4-bit
- SmolLM2-135MHuggingFaceTB0.1B≈0.1 GB at 4-bit
- roberta-baseFacebook AI community0.1B≈0.1 GB at 4-bit
- roberta-base-squad2deepset0.1B≈0.1 GB at 4-bit
- bert-base-uncasedBERT community0.1B≈0.1 GB at 4-bit
- bert-base-casedBERT community0.1B≈0.1 GB at 4-bit
- bert-base-chineseBERT community0.1B≈0.1 GB at 4-bit
- BEN2Prama LLC0.1B≈0.1 GB at 4-bit
- distilgpt2DistilBERT community0.1B≈0.1 GB at 4-bit
- vit-base-patch16-224Google0.1B≈0.1 GB at 4-bit
- nsfw_image_detectionFalcons.ai0.1B≈0.1 GB at 4-bit
- distilbert-base-uncasedDistilBERT community0.1B≈0.0 GB at 4-bit
- t5-smallT5 community0.1B≈0.0 GB at 4-bit
- RMBG-1.4BRIA AI0.0B≈0.0 GB at 4-bit
- detr-resnet-50AI at Meta0.0B≈0.0 GB at 4-bit
- table-transformer-structure-recognitionMicrosoft0.0B≈0.0 GB at 4-bit
- table-transformer-detectionMicrosoft0.0B≈0.0 GB at 4-bit
- segformer_b2_clothesmattmdjaga0.0B≈0.0 GB at 4-bit
- segformer-b0-finetuned-ade-512-512NVIDIA0.0B≈0.0 GB at 4-bit
- Ornith-1.0-9BOrnith0.0B≈0.0 GB at 4-bit
- AutoGLM-Phone-9BZ.ai0.0B≈0.0 GB at 4-bit