Models that run on 8 GB
557 models with published weights that fit in 8 GB — a phone, a base iPad, an Air. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 7 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 557 in all, a hundred to a page; this is page 1 of 6.
The memory of a phone, a base iPad or an entry-level laptop. What fits is small: models of a few billion parameters, quick and cheap to run, good at summarising, classifying and simple extraction, and out of their depth on long reasoning. This is also where on-device makes the most sense, because the alternative is a network round trip for something that takes a moment.
- bloom-7b1BigScience Workshop7.1B≈4.6 GB at 4-bit
- llava-1.5-7b-hfLlava Hugging Face7.1B≈4.6 GB at 4-bit
- DeciLM-7BDeci AI7.0B≈4.6 GB at 4-bit
- chameleon-7bfacebook7.0B≈4.6 GB at 4-bit
- ALLaM-7B-Instruct-previewHUMAIN7.0B≈4.6 GB at 4-bit
- AlphaMonarch-7Bmlabonne7.0B≈4.5 GB at 4-bit
- BELLE-7B-2MBelleGroup7.0B≈4.5 GB at 4-bit
- BioMistral-7BBioMistral7.0B≈4.5 GB at 4-bit
- Chinese-Llama-2-7bLinkSoul7.0B≈4.5 GB at 4-bit
- CodeLlama-7b-Python-hfCode Llama7.0B≈4.5 GB at 4-bit
- Cosmos-1.0-Diffusion-7B-Text2WorldNVIDIA7.0B≈4.5 GB at 4-bit
- DeepSeek-R1-Distill-QWEN-7BDeepSeek7.0B≈4.5 GB at 4-bit3 also selling it hosted
- Dream-v0-Instruct-7BDream-org7.0B≈4.5 GB at 4-bit
- FastVLM-7BApple7.0B≈4.5 GB at 4-bit
- GritLM-7BGritLM7.0B≈4.5 GB at 4-bit
- Hy-MT2-7BTencent Hunyuan7.0B≈4.5 GB at 4-bit1 also selling it hosted
- Janus-Pro-7BDeepSeek7.0B≈4.5 GB at 4-bit1 also selling it hosted
- K2-Horizon-7BInstitute of Foundation Models7.0B≈4.5 GB at 4-bit
- Lily-Cybersecurity-7B-v0.2segolilylabs7.0B≈4.5 GB at 4-bit
- Llama-2-7bMeta Llama7.0B≈4.5 GB at 4-bit
- Llama-2-7b-chatMeta Llama7.0B≈4.5 GB at 4-bit
- Llama2-Chinese-7b-ChatFlagAlpha7.0B≈4.5 GB at 4-bit
- MiMo-7B-RLXiaomiMiMo7.0B≈4.5 GB at 4-bit
- MiMo-Audio-7B-InstructXiaomiMiMo7.0B≈4.5 GB at 4-bit
- MiMo-VL-7B-RLXiaomiMiMo7.0B≈4.5 GB at 4-bit
- Mistral-7B-Instruct-v0.1Mistral AI7.0B≈4.5 GB at 4-bit3 also selling it hosted
- Mistral-7B-Instruct-v0.2Mistral AI7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Mistral-7B-Instruct-v0.3Mistral AI7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Mistral-7B-OpenOrcaOpenOrca7.0B≈4.5 GB at 4-bit
- Mistral-Trismegistus-7Bteknium7.0B≈4.5 GB at 4-bit
- Molmo-7B-O-0924Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit
- NeuralBeagle14-7Bmlabonne7.0B≈4.5 GB at 4-bit
- NeuralHermes-2.5-Mistral-7Bmlabonne7.0B≈4.5 GB at 4-bit
- Nxcode-CQ-7B-orpoNTQAI7.0B≈4.5 GB at 4-bit
- OLMoE-1B-7B-0924Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit
- OlympicCoder-7Bopen-r17.0B≈4.5 GB at 4-bit
- OpenHermes-2-Mistral-7Bteknium7.0B≈4.5 GB at 4-bit1 also selling it hosted
- Openhermes2.5 Mistral 7Bteknium7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Orca-2-7bMicrosoft7.0B≈4.5 GB at 4-bit
- Qwen 2 7B InstructAlibaba7.0B≈4.5 GB at 4-bit4 also selling it hosted
- Qwen2-7BAlibaba7.0B≈4.5 GB at 4-bit
- Qwen2.5-Coder-7BAlibaba7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Qwen2.5-Coder-7B-InstructAlibaba7.0B≈4.5 GB at 4-bit4 also selling it hosted
- Seed-X-PPO-7BByteDance Seed7.0B≈4.5 GB at 4-bit
- TAIDE-LX-7B-Chattaide7.0B≈4.5 GB at 4-bit
- UI-TARS-7B-SFTByteDance Seed7.0B≈4.5 GB at 4-bit
- Vistral-7B-ChatViet-Mistral7.0B≈4.5 GB at 4-bit
- WizardLM-7B-UncensoredQuixi AI7.0B≈4.5 GB at 4-bit
- chinese-alpaca-2-7bJoint Laboratory of HIT and iFLYTEK Research (HFL)7.0B≈4.5 GB at 4-bit
- codegemma-7b-itGoogle7.0B≈4.5 GB at 4-bit1 also selling it hosted
- deepseek-coder-7b-instruct-v1.5deepseek-ai7.0B≈4.5 GB at 4-bit1 also selling it hosted
- deepseek-llm-7b-basedeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-llm-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-math-7b-instructdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-vl-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- e5-mistral-7b-instructintfloat7.0B≈4.5 GB at 4-bit
- gemma-1.1-7b-itGoogle7.0B≈4.5 GB at 4-bit1 also selling it hosted
- gte-Qwen2-7B-instructAlibaba-NLP7.0B≈4.5 GB at 4-bit
- internlm-xcomposer2d5-7binternlm7.0B≈4.5 GB at 4-bit
- llava-v1.6-mistral-7b-hfLlava Hugging Face7.0B≈4.5 GB at 4-bit
- mixtral-7b-8expertDiscoResearch7.0B≈4.5 GB at 4-bit
- olmOCR-2-7B-1025Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit1 also selling it hosted
- olmOCR-7B-0725-FP8Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit1 also selling it hosted
- olmOCR-7B-0825Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit1 also selling it hosted
- open-calm-7bCyberAgent7.0B≈4.5 GB at 4-bit
- rwkv-4-pile-7bBlinkDL7.0B≈4.5 GB at 4-bit
- speed-embedding-7b-instructHaon-Chen7.0B≈4.5 GB at 4-bit
- stablelm-base-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- stablelm-tuned-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- vicuna-7b-v1.5Large Model Systems Organization7.0B≈4.5 GB at 4-bit
- xgen-7b-8k-baseSalesforce7.0B≈4.5 GB at 4-bit
- granite-4.0-h-tinyIBM Granite6.9B≈4.5 GB at 4-bit
- NuminaMath-7B-TIRProject-Numina6.9B≈4.5 GB at 4-bit
- OLMo-7BAi26.9B≈4.5 GB at 4-bit
- Magicoder-S-DS-6.7BIntellligent Software Engineering (iSE)6.7B≈4.4 GB at 4-bit
- deepseek-coder-6.7b-instructDeepSeek6.7B≈4.4 GB at 4-bit
- meditron-7bEPFL LLM Team6.7B≈4.4 GB at 4-bit
- CodeLlama-7b-Instruct-hfCode Llama6.7B≈4.4 GB at 4-bit
- CodeLlama-7b-hfCode Llama6.7B≈4.4 GB at 4-bit
- sqlcoder-7b-2Defog.ai6.7B≈4.4 GB at 4-bit
- Llama-2-7b-chat-hfMeta Llama6.7B≈4.4 GB at 4-bit
- Llama-2-7b-hfMeta Llama6.7B≈4.4 GB at 4-bit
- llama-7bhuggyllama6.7B≈4.4 GB at 4-bit
- llama2_7b_chat_uncensoredgeorgesung6.7B≈4.4 GB at 4-bit
- LlamaGuard-7bMeta Llama6.7B≈4.4 GB at 4-bit1 also selling it hosted
- dinov3-vit7b16-pretrain-lvd1689mfacebook6.7B≈4.4 GB at 4-bit
- YuE-s1-7B-anneal-en-cotMultimodal Art Projection6.2B≈4.0 GB at 4-bit
- Z-ImageTongyi-MAI6.2B≈4.0 GB at 4-bit
- Z-Image-TurboTongyi-MAI6.2B≈4.0 GB at 4-bit
- Yi-6B01-ai6.1B≈3.9 GB at 4-bit1 also selling it hosted
- Yi-6B-200K01-ai6.0B≈3.9 GB at 4-bit
- gpt-j-6bEleutherAI6.0B≈3.9 GB at 4-bit
- pygmalion-6bPygmalion6.0B≈3.9 GB at 4-bit
- Chroma-4BFlashLabs5.9B≈3.8 GB at 4-bit
- higgs-tts-2-3b-baseBoson AI5.8B≈3.8 GB at 4-bit
- DeciLM-6bDeci AI5.7B≈3.7 GB at 4-bit
- CogVideoX-5bZ.ai5.6B≈3.6 GB at 4-bit
- Qwen2.5-Omni-3BAlibaba5.5B≈3.6 GB at 4-bit
- chandra-ocr-2Datalab5.3B≈3.4 GB at 4-bit
- Gemma-4-E2B-itGoogle5.1B≈3.3 GB at 4-bit1 also selling it hosted