Models that run on 8 GB
557 models with published weights that fit in 8 GB — a phone, a base iPad, an Air. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 7 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 557 in all, a hundred to a page; this is page 3 of 6.
The memory of a phone, a base iPad or an entry-level laptop. What fits is small: models of a few billion parameters, quick and cheap to run, good at summarising, classifying and simple extraction, and out of their depth on long reasoning. This is also where on-device makes the most sense, because the alternative is a network round trip for something that takes a moment.
- open_llama_3b_v2OpenLM Research3.0B≈2.0 GB at 4-bit
- orca_mini_3bpankajmathur3.0B≈2.0 GB at 4-bit
- orpheus-3b-0.1-pretrainedCanopy Labs3.0B≈2.0 GB at 4-bit
- paligemma2-3b-pt-224Google3.0B≈2.0 GB at 4-bit
- proxy-lite-3bconvergence-ai3.0B≈2.0 GB at 4-bit
- replit-code-v1-3bReplit3.0B≈2.0 GB at 4-bit
- replit-code-v1_5-3bReplit3.0B≈2.0 GB at 4-bit
- stablecode-completion-alpha-3b-4kStability AI3.0B≈2.0 GB at 4-bit
- stablecode-instruct-alpha-3bStability AI3.0B≈2.0 GB at 4-bit
- stablelm-3b-4e1tStability AI3.0B≈2.0 GB at 4-bit
- tada-3b-mlHume AI3.0B≈2.0 GB at 4-bit
- madlad400-3b-mtGoogle2.9B≈1.9 GB at 4-bit
- paligemma-3b-pt-224Google2.9B≈1.9 GB at 4-bit
- Anima-2.9BGazingstars1232.9B≈1.9 GB at 4-bit
- stable-code-3bStability AI2.8B≈1.8 GB at 4-bit1 also selling it hosted
- stable-code-instruct-3bStability AI2.8B≈1.8 GB at 4-bit
- stablelm-zephyr-3bStability AI2.8B≈1.8 GB at 4-bit
- dolphin-2_6-phi-2Dolphin2.8B≈1.8 GB at 4-bit
- phi-2Microsoft2.8B≈1.8 GB at 4-bit
- Solidity-LLMChainGPT2.8B≈1.8 GB at 4-bit
- gpt-neo-2.7BEleutherAI2.7B≈1.8 GB at 4-bit
- VibeVoice-1.5BMicrosoft2.7B≈1.8 GB at 4-bit
- gemma-2-2bGoogle2.6B≈1.7 GB at 4-bit
- gemma-2-2b-itGoogle2.6B≈1.7 GB at 4-bit
- gemma-2-2b-jpn-itGoogle2.6B≈1.7 GB at 4-bit
- Lumina-Image-2.0Alpha-VLLM2.6B≈1.7 GB at 4-bit
- LFM2-2.6B-TranscriptLiquid AI2.6B≈1.7 GB at 4-bit
- Ouro-2.6B-ThinkingByteDance2.6B≈1.7 GB at 4-bit
- KolorsKolors Team, Kuaishou Technology2.6B≈1.7 GB at 4-bit
- Ovis2.5-2BATH-MaaS2.6B≈1.7 GB at 4-bit
- LFM2-2.6BLiquid AI2.6B≈1.7 GB at 4-bit
- LFM2-2.6B-ExpLiquid AI2.6B≈1.7 GB at 4-bit
- stable-diffusion-xl-1.0-inpainting-0.1🧨Diffusers2.6B≈1.7 GB at 4-bit
- Illustrious-xl-early-release-v0OnomaAI2.6B≈1.7 GB at 4-bit
- OpenDalleV1.1dataautogpt32.6B≈1.7 GB at 4-bit
- SD XLStability AI2.6B≈1.7 GB at 4-bit2 also selling it hosted
- animagine-xl-2.0Linaqruf2.6B≈1.7 GB at 4-bit
- animagine-xl-3.0Cagliostro Labs2.6B≈1.7 GB at 4-bit
- animagine-xl-3.1Cagliostro Labs2.6B≈1.7 GB at 4-bit
- animagine-xl-4.0Cagliostro Labs2.6B≈1.7 GB at 4-bit
- dpo-sdxl-text2image-v1mhdang2.6B≈1.7 GB at 4-bit
- playground-v2-1024px-aestheticPlayground2.6B≈1.7 GB at 4-bit1 also selling it hosted
- playground-v2.5-1024px-aestheticPlayground2.6B≈1.7 GB at 4-bit1 also selling it hosted
- sdxl-flashStable Diffusion Community (Unofficial, Non-profit)2.6B≈1.7 GB at 4-bit
- stable-diffusion-xl-base-0.9Stability AI2.6B≈1.7 GB at 4-bit
- MiniCPM5-2BOpenBMB2.5B≈1.6 GB at 4-bit
- Octopus-v2Nexa AI2.5B≈1.6 GB at 4-bit
- gemma-2bGoogle2.5B≈1.6 GB at 4-bit
- gemma-2b-itGoogle2.5B≈1.6 GB at 4-bit1 also selling it hosted
- BidirLM-Omni-2.5B-EmbeddingBidirLM2.5B≈1.6 GB at 4-bit
- canary-qwen-2.5bNVIDIA2.5B≈1.6 GB at 4-bit
- Cosmos-Reason2-2BNVIDIA2.4B≈1.6 GB at 4-bit
- EXAONE-3.5-2.4B-InstructLGAI-EXAONE2.4B≈1.6 GB at 4-bit
- seamless-m4t-v2-largeAI at Meta2.3B≈1.5 GB at 4-bit
- VoxCPM2OpenBMB2.3B≈1.5 GB at 4-bit
- stable-diffusion-xl-refiner-0.9Stability AI2.3B≈1.5 GB at 4-bit
- stable-diffusion-xl-refiner-1.0Stability AI2.3B≈1.5 GB at 4-bit
- GLM-ASR-Nano-2512Z.ai2.3B≈1.5 GB at 4-bit
- SmolVLM2-2.2B-InstructHugging Face Smol Models Research2.2B≈1.5 GB at 4-bit
- Marlin-2BNemo Station2.2B≈1.4 GB at 4-bit
- Qwen2-VL-2B-InstructAlibaba2.2B≈1.4 GB at 4-bit1 also selling it hosted
- tada-1bHume AI2.2B≈1.4 GB at 4-bit
- FLUX.1-dev-Controlnet-Inpainting-Betaalimama-creative2.1B≈1.4 GB at 4-bit
- FLUX.1-dev-ControlNet-Union-Pro-2.0Shakker Labs2.1B≈1.4 GB at 4-bit
- Qwen3-VL-2B-InstructAlibaba2.1B≈1.4 GB at 4-bit
- Qwen3-VL-Embedding-2BAlibaba2.1B≈1.4 GB at 4-bit
- Janus-1.3BDeepSeek2.1B≈1.4 GB at 4-bit
- stable-diffusion-3-medium-diffusersStability AI2.1B≈1.4 GB at 4-bit
- cohere-transcribe-arabic-07-2026Cohere Labs2.1B≈1.3 GB at 4-bit
- SoulX-Podcast-1.7BSoul-AILab2.1B≈1.3 GB at 4-bit
- Qwen3-1.7BAlibaba2.0B≈1.3 GB at 4-bit1 also selling it hosted
- Isaac 0.2 2B PreviewPerceptron2.0B≈1.3 GB at 4-bit1 also selling it hosted
- MOSS-Transcribe-preview-2BOpenMOSS-Team2.0B≈1.3 GB at 4-bit
- MiniCPM-2B-sft-fp32OpenBMB2.0B≈1.3 GB at 4-bit
- Qwen3.5-2BAlibaba2.0B≈1.3 GB at 4-bit2 also selling it hosted
- gemma-1.1-2b-itGoogle2.0B≈1.3 GB at 4-bit
- granite-speech-4.1-2bIBM watsonx.ai2.0B≈1.3 GB at 4-bit
- granite-speech-4.1-2b-narIBM watsonx.ai2.0B≈1.3 GB at 4-bit
- helium-1-preview-2bKyutai2.0B≈1.3 GB at 4-bit
- Youtu-LLM-2BTencent Hunyuan2.0B≈1.3 GB at 4-bit
- moondream2vikhyatk1.9B≈1.3 GB at 4-bit
- LTX-VideoLTX.io1.9B≈1.3 GB at 4-bit
- Dia2-2BNari Labs1.9B≈1.2 GB at 4-bit
- Qwen3-TTS-12Hz-1.7B-CustomVoiceAlibaba1.9B≈1.2 GB at 4-bit
- Qwen3-TTS-12Hz-1.7B-VoiceDesignAlibaba1.9B≈1.2 GB at 4-bit
- Hy-MT2-1.8BTencent Hunyuan1.8B≈1.2 GB at 4-bit1 also selling it hosted
- Flux.1-dev-Controlnet-UpscalerJasper.ai1.8B≈1.2 GB at 4-bit
- DeepScaleR-1.5B-PreviewAgentica1.8B≈1.2 GB at 4-bit
- DeepSeek-R1-Distill-Qwen-1.5BDeepSeek1.8B≈1.2 GB at 4-bit3 also selling it hosted
- Nemotron-Research-Reasoning-Qwen-1.5BNVIDIA1.8B≈1.2 GB at 4-bit
- VibeThinker-1.5BWeiboAI1.8B≈1.2 GB at 4-bit
- gte-Qwen2-1.5B-instructAlibaba-NLP1.8B≈1.2 GB at 4-bit
- Qwen3.6-27B-DFlashZ Lab1.7B≈1.1 GB at 4-bit
- BidirLM-1.7B-EmbeddingBidirLM1.7B≈1.1 GB at 4-bit
- F2LLM-1.7Bcodefuse-ai1.7B≈1.1 GB at 4-bit
- F2LLM-v2-1.7Bcodefuse-ai1.7B≈1.1 GB at 4-bit
- Qwen3-ASR-1.7BAlibaba1.7B≈1.1 GB at 4-bit1 also selling it hosted
- SmolLM-1.7BHuggingFaceTB1.7B≈1.1 GB at 4-bit
- SmolLM2-1.7BHuggingFaceTB1.7B≈1.1 GB at 4-bit
- CogVideoX-2bZ.ai1.7B≈1.1 GB at 4-bit