Models that run on 32 GB
986 models with published weights that fit in 32 GB — a well-specified laptop. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 33 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 986 in all, a hundred to a page; this is page 9 of 10.
Enough for a thirty-billion-parameter model at four-bit with a long context, or a smaller one at higher precision if quality matters more than size. A practical ceiling for a laptop that also has to be a laptop.
- Gemma3-1B-ITLiteRT Community (FKA TFLite)1.0B≈0.7 GB at 4-bit
- Isaac 0.2 1BPerceptron1.0B≈0.7 GB at 4-bit1 also selling it hosted
- Janus-Pro-1BDeepSeek1.0B≈0.7 GB at 4-bit1 also selling it hosted
- MiniCPM5-1B-Claude-Opus-Fable5-ThinkingGnLOLot1.0B≈0.7 GB at 4-bit
- MolmoE-1B-0924Allen Institute for AI (Ai2)1.0B≈0.7 GB at 4-bit
- Nemotron-3-Embed-1B-BF16NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- Nemotron-3-Embed-1B-NVFP4NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- Taiyi-Stable-Diffusion-1B-Chinese-v0.1Fengshenbang-LM1.0B≈0.7 GB at 4-bit
- antares-1bfdtn-ai1.0B≈0.7 GB at 4-bit
- canary-1bNVIDIA1.0B≈0.7 GB at 4-bit
- canary-1b-flashNVIDIA1.0B≈0.7 GB at 4-bit
- csm-1bsesame1.0B≈0.7 GB at 4-bit1 also selling it hosted
- granite-4.0-1b-speechIBM watsonx.ai1.0B≈0.7 GB at 4-bit
- granite-4.0-h-1bIBM Granite1.0B≈0.7 GB at 4-bit
- llama-nemotron-embed-vl-1b-v2NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- llama-nemotron-rerank-vl-1b-v2NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- metavoice-1B-v0.1MetaVoice1.0B≈0.7 GB at 4-bit
- gemma-3-1b-itGoogle1.0B≈0.6 GB at 4-bit
- gemma-3-1b-ptGoogle1.0B≈0.6 GB at 4-bit
- CLIP-ViT-H-14-laion2B-s32B-b79KLAION eV1.0B≈0.6 GB at 4-bit
- canary-1b-v2NVIDIA1.0B≈0.6 GB at 4-bit
- mms-1b-allfacebook1.0B≈0.6 GB at 4-bit
- PaddleOCR-VL-1.5paddlepaddle1.0B≈0.6 GB at 4-bit
- PaddleOCR-VL-1.6paddlepaddle1.0B≈0.6 GB at 4-bit
- MobileLLM-R1-950MAI at Meta0.9B≈0.6 GB at 4-bit
- indic-parler-ttsAI4Bharat0.9B≈0.6 GB at 4-bit
- Qwen3-TTS-12Hz-0.6B-CustomVoiceAlibaba0.9B≈0.6 GB at 4-bit
- PaddleOCR-VL-0.9Bpaddlepaddle0.9B≈0.6 GB at 4-bit1 also selling it hosted
- siglip-so400m-patch14-384Google0.9B≈0.6 GB at 4-bit
- waifu-diffusionhakurei0.9B≈0.6 GB at 4-bit
- LCM_Dreamshaper_v7SimianLuo0.9B≈0.6 GB at 4-bit
- instruct-pix2pixtimbrooks0.9B≈0.6 GB at 4-bit
- Cyberpunk-Anime-DiffusionDGSpitzer0.9B≈0.6 GB at 4-bit
- Dungeons-and-Diffusion0xJustin0.9B≈0.6 GB at 4-bit
- EimisAnimeDiffusion_1.0veimiss0.9B≈0.6 GB at 4-bit
- Ghibli-Diffusionnitrosocke0.9B≈0.6 GB at 4-bit
- GuoFeng3xiaolxl0.9B≈0.6 GB at 4-bit
- Nitro-Diffusionnitrosocke0.9B≈0.6 GB at 4-bit
- Stable_Diffusion_PaperCut_ModelFictiverse0.9B≈0.6 GB at 4-bit
- anything-v5Stable Diffusion API0.9B≈0.6 GB at 4-bit
- dreamlike-anime-1.0dreamlike-art0.9B≈0.6 GB at 4-bit
- dreamlike-diffusion-1.0dreamlike-art0.9B≈0.6 GB at 4-bit
- dreamlike-photoreal-2.0dreamlike-art0.9B≈0.6 GB at 4-bit
- openjourney-v4prompthero0.9B≈0.6 GB at 4-bit
- OvisOCR2ATH-MaaS0.9B≈0.6 GB at 4-bit
- bitnet-b1.58-2B-4TMicrosoft0.8B≈0.6 GB at 4-bit
- gpt2-largeOpenAI community0.8B≈0.5 GB at 4-bit
- VoxCPM1.5OpenBMB0.8B≈0.5 GB at 4-bit
- Qwen3.5-0.8BAlibaba0.8B≈0.5 GB at 4-bit1 also selling it hosted
- t5gemma-2-270m-270mGoogle0.8B≈0.5 GB at 4-bit
- Florence-2-largeMicrosoft0.8B≈0.5 GB at 4-bit
- Florence-2-large-ftMicrosoft0.8B≈0.5 GB at 4-bit
- FastVLM-0.5BApple0.8B≈0.5 GB at 4-bit
- distil-large-v3Whisper Distillation0.8B≈0.5 GB at 4-bit
- distil-large-v2Whisper Distillation0.8B≈0.5 GB at 4-bit
- Qwen3-0.6BAlibaba0.8B≈0.5 GB at 4-bit1 also selling it hosted
- neutts-airNeuphonic0.7B≈0.5 GB at 4-bit
- GOT-OCR2_0StepFun0.7B≈0.5 GB at 4-bit
- parler_tts_mini_v0.1Parler TTS0.6B≈0.4 GB at 4-bit
- NVIDIA Nemotron 3.5 ASR Streaming 0.6BNVIDIA0.6B≈0.4 GB at 4-bit
- Parakeet TDT 0.6B v3NVIDIA0.6B≈0.4 GB at 4-bit
- nemotron-speech-streaming-en-0.6bNVIDIA0.6B≈0.4 GB at 4-bit
- PixArt-XL-2-1024-MSPixArt0.6B≈0.4 GB at 4-bit
- Audio8-TTS-Preview-0.6bAudio80.6B≈0.4 GB at 4-bit1 also selling it hosted
- F2LLM-0.6Bcodefuse-ai0.6B≈0.4 GB at 4-bit
- F2LLM-v2-0.6Bcodefuse-ai0.6B≈0.4 GB at 4-bit
- Qwen3-ASR-0.6BAlibaba0.6B≈0.4 GB at 4-bit1 also selling it hosted
- Qwen3-Embedding-0.6BAlibaba0.6B≈0.4 GB at 4-bit4 also selling it hosted
- Qwen3-ForcedAligner-0.6BAlibaba0.6B≈0.4 GB at 4-bit
- Qwen3-Reranker-0.6BAlibaba0.6B≈0.4 GB at 4-bit1 also selling it hosted
- Qwen3-TTS-12Hz-0.6B-BaseAlibaba0.6B≈0.4 GB at 4-bit
- harrier-oss-v1-0.6bMicrosoft0.6B≈0.4 GB at 4-bit
- parakeet-tdt-0.6b-v2NVIDIA0.6B≈0.4 GB at 4-bit
- Qwen3-0.6B-BaseAlibaba0.6B≈0.4 GB at 4-bit
- musicgen-smallAI at Meta0.6B≈0.4 GB at 4-bit
- w2v-bert-2.0facebook0.6B≈0.4 GB at 4-bit
- jina-embeddings-v3Jina AI0.6B≈0.4 GB at 4-bit
- xlm-roberta-largeFacebook AI community0.6B≈0.4 GB at 4-bit
- GOT-OCR-2.0-hfstepfun-ai0.6B≈0.4 GB at 4-bit
- bge-reranker-largeBAAI0.6B≈0.4 GB at 4-bit
- Multilingual-e5-large-instructintfloat0.6B≈0.4 GB at 4-bit4 also selling it hosted
- bloom-560mBigScience Workshop0.6B≈0.4 GB at 4-bit
- MiraTTSYatharthS0.5B≈0.3 GB at 4-bit
- Fun-CosyVoice3-0.5B-2512QwenAudio0.5B≈0.3 GB at 4-bit
- Qwen1.5-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- VoxCPM-0.5BOpenBMB0.5B≈0.3 GB at 4-bit
- reader-lm-0.5bJina AI0.5B≈0.3 GB at 4-bit
- Qwen2-0.5B-InstructAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2.5-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2.5-0.5B-InstructAlibaba0.5B≈0.3 GB at 4-bit1 also selling it hosted
- nomic-embed-text-v2-moeNomic AI0.5B≈0.3 GB at 4-bit
- LaBSEsentence-transformers0.5B≈0.3 GB at 4-bit
- blip-image-captioning-largeSalesforce0.5B≈0.3 GB at 4-bit
- LFM2.5-VL-450MLiquid AI0.4B≈0.3 GB at 4-bit
- PairRMLLM Blender0.4B≈0.3 GB at 4-bit
- stella_en_400M_v5NovaSearch0.4B≈0.3 GB at 4-bit
- gte-large-en-v1.5Alibaba-NLP0.4B≈0.3 GB at 4-bit
- CLIP-GmP-ViT-L-14zer0int0.4B≈0.3 GB at 4-bit
- circuit-sparsityOpenAI0.4B≈0.3 GB at 4-bit