Models that run on 64 GB
1047 models with published weights that fit in 64 GB — a workstation. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 67 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 1,047 in all, a hundred to a page; this is page 10 of 11.
A workstation, or a well-specified Mac. Seventy-billion-parameter models fit at four-bit, which is the class where local output stops being obviously worse than the hosted models people pay for. Loading takes real time and the machine will be warm.
- anything-v5Stable Diffusion API0.9B≈0.6 GB at 4-bit
- dreamlike-anime-1.0dreamlike-art0.9B≈0.6 GB at 4-bit
- dreamlike-diffusion-1.0dreamlike-art0.9B≈0.6 GB at 4-bit
- dreamlike-photoreal-2.0dreamlike-art0.9B≈0.6 GB at 4-bit
- openjourney-v4prompthero0.9B≈0.6 GB at 4-bit
- OvisOCR2ATH-MaaS0.9B≈0.6 GB at 4-bit
- bitnet-b1.58-2B-4TMicrosoft0.8B≈0.6 GB at 4-bit
- gpt2-largeOpenAI community0.8B≈0.5 GB at 4-bit
- VoxCPM1.5OpenBMB0.8B≈0.5 GB at 4-bit
- Qwen3.5-0.8BAlibaba0.8B≈0.5 GB at 4-bit1 also selling it hosted
- t5gemma-2-270m-270mGoogle0.8B≈0.5 GB at 4-bit
- Florence-2-largeMicrosoft0.8B≈0.5 GB at 4-bit
- Florence-2-large-ftMicrosoft0.8B≈0.5 GB at 4-bit
- FastVLM-0.5BApple0.8B≈0.5 GB at 4-bit
- distil-large-v3Whisper Distillation0.8B≈0.5 GB at 4-bit
- distil-large-v2Whisper Distillation0.8B≈0.5 GB at 4-bit
- Qwen3-0.6BAlibaba0.8B≈0.5 GB at 4-bit1 also selling it hosted
- neutts-airNeuphonic0.7B≈0.5 GB at 4-bit
- GOT-OCR2_0StepFun0.7B≈0.5 GB at 4-bit
- parler_tts_mini_v0.1Parler TTS0.6B≈0.4 GB at 4-bit
- NVIDIA Nemotron 3.5 ASR Streaming 0.6BNVIDIA0.6B≈0.4 GB at 4-bit
- Parakeet TDT 0.6B v3NVIDIA0.6B≈0.4 GB at 4-bit
- nemotron-speech-streaming-en-0.6bNVIDIA0.6B≈0.4 GB at 4-bit
- PixArt-XL-2-1024-MSPixArt0.6B≈0.4 GB at 4-bit
- Audio8-TTS-Preview-0.6bAudio80.6B≈0.4 GB at 4-bit1 also selling it hosted
- F2LLM-0.6Bcodefuse-ai0.6B≈0.4 GB at 4-bit
- F2LLM-v2-0.6Bcodefuse-ai0.6B≈0.4 GB at 4-bit
- Qwen3-ASR-0.6BAlibaba0.6B≈0.4 GB at 4-bit1 also selling it hosted
- Qwen3-Embedding-0.6BAlibaba0.6B≈0.4 GB at 4-bit4 also selling it hosted
- Qwen3-ForcedAligner-0.6BAlibaba0.6B≈0.4 GB at 4-bit
- Qwen3-Reranker-0.6BAlibaba0.6B≈0.4 GB at 4-bit1 also selling it hosted
- Qwen3-TTS-12Hz-0.6B-BaseAlibaba0.6B≈0.4 GB at 4-bit
- harrier-oss-v1-0.6bMicrosoft0.6B≈0.4 GB at 4-bit
- parakeet-tdt-0.6b-v2NVIDIA0.6B≈0.4 GB at 4-bit
- Qwen3-0.6B-BaseAlibaba0.6B≈0.4 GB at 4-bit
- musicgen-smallAI at Meta0.6B≈0.4 GB at 4-bit
- w2v-bert-2.0facebook0.6B≈0.4 GB at 4-bit
- jina-embeddings-v3Jina AI0.6B≈0.4 GB at 4-bit
- xlm-roberta-largeFacebook AI community0.6B≈0.4 GB at 4-bit
- GOT-OCR-2.0-hfstepfun-ai0.6B≈0.4 GB at 4-bit
- bge-reranker-largeBAAI0.6B≈0.4 GB at 4-bit
- Multilingual-e5-large-instructintfloat0.6B≈0.4 GB at 4-bit4 also selling it hosted
- bloom-560mBigScience Workshop0.6B≈0.4 GB at 4-bit
- MiraTTSYatharthS0.5B≈0.3 GB at 4-bit
- Fun-CosyVoice3-0.5B-2512QwenAudio0.5B≈0.3 GB at 4-bit
- Qwen1.5-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- VoxCPM-0.5BOpenBMB0.5B≈0.3 GB at 4-bit
- reader-lm-0.5bJina AI0.5B≈0.3 GB at 4-bit
- Qwen2-0.5B-InstructAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2.5-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2.5-0.5B-InstructAlibaba0.5B≈0.3 GB at 4-bit1 also selling it hosted
- nomic-embed-text-v2-moeNomic AI0.5B≈0.3 GB at 4-bit
- LaBSEsentence-transformers0.5B≈0.3 GB at 4-bit
- blip-image-captioning-largeSalesforce0.5B≈0.3 GB at 4-bit
- LFM2.5-VL-450MLiquid AI0.4B≈0.3 GB at 4-bit
- PairRMLLM Blender0.4B≈0.3 GB at 4-bit
- stella_en_400M_v5NovaSearch0.4B≈0.3 GB at 4-bit
- gte-large-en-v1.5Alibaba-NLP0.4B≈0.3 GB at 4-bit
- CLIP-GmP-ViT-L-14zer0int0.4B≈0.3 GB at 4-bit
- circuit-sparsityOpenAI0.4B≈0.3 GB at 4-bit
- bart-large-cnnAI at Meta0.4B≈0.3 GB at 4-bit
- DolphinByteDance0.4B≈0.3 GB at 4-bit
- ModernBERT-largeAnswer.AI0.4B≈0.3 GB at 4-bit
- blip-vqa-baseSalesforce0.4B≈0.3 GB at 4-bit
- control_v1p_sd15_brightnessLatent Cat0.4B≈0.2 GB at 4-bit
- controlnet_qrcode-control_v1p_sd15DionTimmer0.4B≈0.2 GB at 4-bit
- LFM2-350MLiquid AI0.4B≈0.2 GB at 4-bit
- LFM2.5-350MLiquid AI0.4B≈0.2 GB at 4-bit
- IndicF5AI4Bharat0.4B≈0.2 GB at 4-bit
- nougat-basefacebook0.3B≈0.2 GB at 4-bit
- FLUX.1-Turbo-Alphaalimama-creative0.3B≈0.2 GB at 4-bit
- bert-large-uncased-whole-word-masking-finetuned-squadBERT community0.3B≈0.2 GB at 4-bit
- bge-large-enBAAI0.3B≈0.2 GB at 4-bit1 also selling it hosted
- UAE-Large-V1WhereIsAI0.3B≈0.2 GB at 4-bit
- mxbai-embed-large-v1Mixedbread0.3B≈0.2 GB at 4-bit
- trocr-base-handwrittenMicrosoft0.3B≈0.2 GB at 4-bit
- bge-large-zhBAAI0.3B≈0.2 GB at 4-bit
- text2vec-large-chineseGanymedeNil0.3B≈0.2 GB at 4-bit
- VoiceCraftpyp10.3B≈0.2 GB at 4-bit
- VibeVoice-ASR-BitNetMicrosoft0.3B≈0.2 GB at 4-bit
- wav2vec2-lg-xlsr-en-speech-emotion-recognitionehcalabres0.3B≈0.2 GB at 4-bit
- wav2vec2-large-xlsr-53-englishjonatasgrosman0.3B≈0.2 GB at 4-bit
- gte-multilingual-baseAlibaba-NLP0.3B≈0.2 GB at 4-bit
- dinov3-vitl16-pretrain-lvd1689mAI at Meta0.3B≈0.2 GB at 4-bit
- xlm-roberta-baseFacebook AI community0.3B≈0.2 GB at 4-bit
- mdeberta-v3-base-squad2timpal0l0.3B≈0.2 GB at 4-bit
- multilingual-e5-baseintfloat0.3B≈0.2 GB at 4-bit
- paraphrase-multilingual-mpnet-base-v2sentence-transformers0.3B≈0.2 GB at 4-bit
- functiongemma-270m-itGoogle0.3B≈0.2 GB at 4-bit
- gemma-3-270mGoogle0.3B≈0.2 GB at 4-bit
- gemma-3-270m-itGoogle0.3B≈0.2 GB at 4-bit
- granite-docling-258MIBM Granite0.3B≈0.2 GB at 4-bit
- NeoBERTChandar Research Lab0.2B≈0.2 GB at 4-bit
- whisper-smallOpenAI0.2B≈0.2 GB at 4-bit
- Florence-2-baseMicrosoft0.2B≈0.2 GB at 4-bit
- t5-baseT5 community0.2B≈0.1 GB at 4-bit
- chatgpt_paraphraser_on_T5_baseHumarin0.2B≈0.1 GB at 4-bit
- jina-clip-v1Jina AI0.2B≈0.1 GB at 4-bit
- BiRefNetZhengPeng70.2B≈0.1 GB at 4-bit