Models that run on 8 GB
557 models with published weights that fit in 8 GB — a phone, a base iPad, an Air. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 7 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 557 in all, a hundred to a page; this is page 6 of 6.
The memory of a phone, a base iPad or an entry-level laptop. What fits is small: models of a few billion parameters, quick and cheap to run, good at summarising, classifying and simple extraction, and out of their depth on long reasoning. This is also where on-device makes the most sense, because the alternative is a network round trip for something that takes a moment.
- gemma-3-270mGoogle0.3B≈0.2 GB at 4-bit
- gemma-3-270m-itGoogle0.3B≈0.2 GB at 4-bit
- granite-docling-258MIBM Granite0.3B≈0.2 GB at 4-bit
- NeoBERTChandar Research Lab0.2B≈0.2 GB at 4-bit
- whisper-smallOpenAI0.2B≈0.2 GB at 4-bit
- Florence-2-baseMicrosoft0.2B≈0.2 GB at 4-bit
- t5-baseT5 community0.2B≈0.1 GB at 4-bit
- chatgpt_paraphraser_on_T5_baseHumarin0.2B≈0.1 GB at 4-bit
- jina-clip-v1Jina AI0.2B≈0.1 GB at 4-bit
- BiRefNetZhengPeng70.2B≈0.1 GB at 4-bit
- RMBG-2.0BRIA AI0.2B≈0.1 GB at 4-bit
- bert-base-multilingual-casedBERT community0.2B≈0.1 GB at 4-bit
- Audio8-TTS-Preview-0.1bEdge00.2B≈0.1 GB at 4-bit
- jina-embeddings-v2-base-zhJina AI0.2B≈0.1 GB at 4-bit
- gpt-neo-125mEleutherAI0.2B≈0.1 GB at 4-bit
- ModernBERT-baseAnswer.AI0.1B≈0.1 GB at 4-bit
- gte-modernbert-baseAlibaba-NLP0.1B≈0.1 GB at 4-bit
- modernbert-embed-baseNomic AI0.1B≈0.1 GB at 4-bit
- bart-basefacebook0.1B≈0.1 GB at 4-bit
- jina-embeddings-v2-base-enJina AI0.1B≈0.1 GB at 4-bit
- MagicPrompt-Stable-DiffusionGustavosta0.1B≈0.1 GB at 4-bit
- gpt2OpenAI community0.1B≈0.1 GB at 4-bit
- nomic-embed-text-v1Nomic AI0.1B≈0.1 GB at 4-bit
- distilbert-base-multilingual-caseddistilbert0.1B≈0.1 GB at 4-bit
- distiluse-base-multilingual-cased-v2sentence-transformers0.1B≈0.1 GB at 4-bit
- SmolLM2-135MHuggingFaceTB0.1B≈0.1 GB at 4-bit
- roberta-baseFacebook AI community0.1B≈0.1 GB at 4-bit
- roberta-base-squad2deepset0.1B≈0.1 GB at 4-bit
- openjourneyprompthero0.1B≈0.1 GB at 4-bit
- paraphrase-multilingual-MiniLM-L12-v2sentence-transformers0.1B≈0.1 GB at 4-bit
- bert-base-uncasedBERT community0.1B≈0.1 GB at 4-bit
- pubmedbert-base-embeddingsNeuML0.1B≈0.1 GB at 4-bit
- bert-base-casedBERT community0.1B≈0.1 GB at 4-bit
- bert-base-chineseBERT community0.1B≈0.1 GB at 4-bit
- BEN2Prama LLC0.1B≈0.1 GB at 4-bit
- wav2vec2-base-960hAI at Meta0.1B≈0.1 GB at 4-bit
- nomic-embed-vision-v1.5Nomic AI0.1B≈0.1 GB at 4-bit
- distilgpt2DistilBERT community0.1B≈0.1 GB at 4-bit
- ast-finetuned-audioset-10-10-0.4593Massachusetts Institute of Technology0.1B≈0.1 GB at 4-bit
- dinov2-basefacebook0.1B≈0.1 GB at 4-bit
- vit-base-patch16-224Google0.1B≈0.1 GB at 4-bit
- vit-base-patch16-224-in21kGoogle0.1B≈0.1 GB at 4-bit
- nsfw_image_detectionFalcons.ai0.1B≈0.1 GB at 4-bit
- Soprano-1.1-80Mekwek0.1B≈0.1 GB at 4-bit
- distilbert-base-uncasedDistilBERT community0.1B≈0.0 GB at 4-bit
- t5-smallT5 community0.1B≈0.0 GB at 4-bit
- RMBG-1.4BRIA AI0.0B≈0.0 GB at 4-bit
- detr-resnet-50AI at Meta0.0B≈0.0 GB at 4-bit
- whisper-tinyOpenAI0.0B≈0.0 GB at 4-bit
- gte-smallthenlper0.0B≈0.0 GB at 4-bit
- table-transformer-structure-recognitionMicrosoft0.0B≈0.0 GB at 4-bit
- table-transformer-detectionMicrosoft0.0B≈0.0 GB at 4-bit
- segformer_b2_clothesmattmdjaga0.0B≈0.0 GB at 4-bit
- dinov3-vits16-pretrain-lvd1689mfacebook0.0B≈0.0 GB at 4-bit
- segformer-b0-finetuned-ade-512-512NVIDIA0.0B≈0.0 GB at 4-bit
- Ornith-1.0-9BOrnith0.0B≈0.0 GB at 4-bit
- AutoGLM-Phone-9BZ.ai0.0B≈0.0 GB at 4-bit