Conversation models that run on 8 GB
257 models with published weights that fit in 8 GB — a phone, a base iPad, an Air. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 7 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 257 in all, a hundred to a page; this is page 1 of 3.
The memory of a phone, a base iPad or an entry-level laptop. What fits is small: models of a few billion parameters, quick and cheap to run, good at summarising, classifying and simple extraction, and out of their depth on long reasoning. This is also where on-device makes the most sense, because the alternative is a network round trip for something that takes a moment.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- bloom-7b1BigScience Workshop7.1B≈4.6 GB at 4-bit
- llava-1.5-7b-hfLlava Hugging Face7.1B≈4.6 GB at 4-bit
- DeciLM-7BDeci AI7.0B≈4.6 GB at 4-bit
- chameleon-7bfacebook7.0B≈4.6 GB at 4-bit
- ALLaM-7B-Instruct-previewHUMAIN7.0B≈4.6 GB at 4-bit
- AlphaMonarch-7Bmlabonne7.0B≈4.5 GB at 4-bit
- BELLE-7B-2MBelleGroup7.0B≈4.5 GB at 4-bit
- BioMistral-7BBioMistral7.0B≈4.5 GB at 4-bit
- Chinese-Llama-2-7bLinkSoul7.0B≈4.5 GB at 4-bit
- DeepSeek-R1-Distill-QWEN-7BDeepSeek7.0B≈4.5 GB at 4-bit3 also selling it hosted
- Dream-v0-Instruct-7BDream-org7.0B≈4.5 GB at 4-bit
- FastVLM-7BApple7.0B≈4.5 GB at 4-bit
- K2-Horizon-7BInstitute of Foundation Models7.0B≈4.5 GB at 4-bit
- Lily-Cybersecurity-7B-v0.2segolilylabs7.0B≈4.5 GB at 4-bit
- Llama-2-7bMeta Llama7.0B≈4.5 GB at 4-bit
- Llama-2-7b-chatMeta Llama7.0B≈4.5 GB at 4-bit
- Llama2-Chinese-7b-ChatFlagAlpha7.0B≈4.5 GB at 4-bit
- MiMo-7B-RLXiaomiMiMo7.0B≈4.5 GB at 4-bit
- MiMo-VL-7B-RLXiaomiMiMo7.0B≈4.5 GB at 4-bit
- Mistral-7B-Instruct-v0.1Mistral AI7.0B≈4.5 GB at 4-bit3 also selling it hosted
- Mistral-7B-Instruct-v0.2Mistral AI7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Mistral-7B-OpenOrcaOpenOrca7.0B≈4.5 GB at 4-bit
- Mistral-Trismegistus-7Bteknium7.0B≈4.5 GB at 4-bit
- Molmo-7B-O-0924Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit
- NeuralBeagle14-7Bmlabonne7.0B≈4.5 GB at 4-bit
- NeuralHermes-2.5-Mistral-7Bmlabonne7.0B≈4.5 GB at 4-bit
- OLMoE-1B-7B-0924Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit
- OpenHermes-2-Mistral-7Bteknium7.0B≈4.5 GB at 4-bit1 also selling it hosted
- Openhermes2.5 Mistral 7Bteknium7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Orca-2-7bMicrosoft7.0B≈4.5 GB at 4-bit
- Qwen 2 7B InstructAlibaba7.0B≈4.5 GB at 4-bit4 also selling it hosted
- Qwen2-7BAlibaba7.0B≈4.5 GB at 4-bit
- TAIDE-LX-7B-Chattaide7.0B≈4.5 GB at 4-bit
- UI-TARS-7B-SFTByteDance Seed7.0B≈4.5 GB at 4-bit
- Vistral-7B-ChatViet-Mistral7.0B≈4.5 GB at 4-bit
- WizardLM-7B-UncensoredQuixi AI7.0B≈4.5 GB at 4-bit
- chinese-alpaca-2-7bJoint Laboratory of HIT and iFLYTEK Research (HFL)7.0B≈4.5 GB at 4-bit
- deepseek-llm-7b-basedeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-llm-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-math-7b-instructdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-vl-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- gemma-1.1-7b-itGoogle7.0B≈4.5 GB at 4-bit1 also selling it hosted
- internlm-xcomposer2d5-7binternlm7.0B≈4.5 GB at 4-bit
- mixtral-7b-8expertDiscoResearch7.0B≈4.5 GB at 4-bit
- open-calm-7bCyberAgent7.0B≈4.5 GB at 4-bit
- rwkv-4-pile-7bBlinkDL7.0B≈4.5 GB at 4-bit
- stablelm-base-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- stablelm-tuned-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- vicuna-7b-v1.5Large Model Systems Organization7.0B≈4.5 GB at 4-bit
- xgen-7b-8k-baseSalesforce7.0B≈4.5 GB at 4-bit
- granite-4.0-h-tinyIBM Granite6.9B≈4.5 GB at 4-bit
- NuminaMath-7B-TIRProject-Numina6.9B≈4.5 GB at 4-bit
- OLMo-7BAi26.9B≈4.5 GB at 4-bit
- meditron-7bEPFL LLM Team6.7B≈4.4 GB at 4-bit
- Llama-2-7b-chat-hfMeta Llama6.7B≈4.4 GB at 4-bit
- Llama-2-7b-hfMeta Llama6.7B≈4.4 GB at 4-bit
- llama-7bhuggyllama6.7B≈4.4 GB at 4-bit
- llama2_7b_chat_uncensoredgeorgesung6.7B≈4.4 GB at 4-bit
- LlamaGuard-7bMeta Llama6.7B≈4.4 GB at 4-bit1 also selling it hosted
- YuE-s1-7B-anneal-en-cotMultimodal Art Projection6.2B≈4.0 GB at 4-bit
- Yi-6B01-ai6.1B≈3.9 GB at 4-bit1 also selling it hosted
- Yi-6B-200K01-ai6.0B≈3.9 GB at 4-bit
- gpt-j-6bEleutherAI6.0B≈3.9 GB at 4-bit
- pygmalion-6bPygmalion6.0B≈3.9 GB at 4-bit
- DeciLM-6bDeci AI5.7B≈3.7 GB at 4-bit
- Mage-VLMicrosoft4.7B≈3.1 GB at 4-bit
- Qwen-Drive-1.0-4BAlibaba4.5B≈3.0 GB at 4-bit
- Llama-3.1-Minitron-4B-Width-BaseNVIDIA4.5B≈2.9 GB at 4-bit
- medgemma-1.5-4b-itGoogle4.3B≈2.8 GB at 4-bit
- medgemma-4b-itGoogle4.3B≈2.8 GB at 4-bit
- NeoHorse-1-4BTokenRhythm4.2B≈2.7 GB at 4-bit
- Nanbeige4.2-3BNanbeige LLM Lab4.2B≈2.7 GB at 4-bit
- Phi-3-vision-128k-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Phi-3.5-vision-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Spark-X2.5-4BSparkLLM4.1B≈2.7 GB at 4-bit
- MiniCPM-V-4OpenBMB4.1B≈2.6 GB at 4-bit
- AgentCPM-ExploreOpenBMB4.0B≈2.6 GB at 4-bit
- Jan-nano-128kMenlo Research4.0B≈2.6 GB at 4-bit
- Jan-v1-4BJan4.0B≈2.6 GB at 4-bit
- Mellum-4b-baseJetBrains4.0B≈2.6 GB at 4-bit
- GELab-Zero-4B-previewstepfun-ai4.0B≈2.6 GB at 4-bit
- Gemma-3-Gaia-PT-BR-4b-itCEIA-UFG4.0B≈2.6 GB at 4-bit
- LocoTrainer-4BLocoreMind4.0B≈2.6 GB at 4-bit
- MiniCPM3-4BOpenBMB4.0B≈2.6 GB at 4-bit
- Nemotron-Mini-4B-InstructNVIDIA4.0B≈2.6 GB at 4-bit
- Qwen3-4b-Z-Image-Engineer-V4BennyDaBall4.0B≈2.6 GB at 4-bit
- Qwen3-VL-4B-InstructAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- R-4BYannQi4.0B≈2.6 GB at 4-bit
- Youtu-VL-4B-InstructTencent Hunyuan4.0B≈2.6 GB at 4-bit
- gemma-3-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- medgemma-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- t5gemma-2-4b-4bGoogle4.0B≈2.6 GB at 4-bit
- Nanbeige4.1-3BNanbeige LLM Lab3.9B≈2.6 GB at 4-bit
- Phi-4-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Phi-3-mini-128k-instructMicrosoft3.8B≈2.5 GB at 4-bit2 also selling it hosted
- Phi-3-mini-4k-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Phi-3.5-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Hy-Embodied-0.5Tencent Hunyuan3.8B≈2.5 GB at 4-bit
- blip2-opt-2.7bSalesforce3.7B≈2.4 GB at 4-bit
- HyperCLOVAX-SEED-Vision-Instruct-3BHyperCLOVA X3.7B≈2.4 GB at 4-bit