Models that run on 96 GB
1115 models with published weights that fit in 96 GB — a MacBook Pro with 96. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 102 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 1,115 in all, a hundred to a page; this is page 11 of 12.
An Apple configuration for people who intend to run models rather than occasionally try one. Comfortably holds the seventy-billion class with a very long context, or a mixture-of-experts model whose active parameters are few but whose weights are all resident.
- harrier-oss-v1-0.6bMicrosoft0.6B≈0.4 GB at 4-bit
- parakeet-tdt-0.6b-v2NVIDIA0.6B≈0.4 GB at 4-bit
- Qwen3-0.6B-BaseAlibaba0.6B≈0.4 GB at 4-bit
- musicgen-smallAI at Meta0.6B≈0.4 GB at 4-bit
- w2v-bert-2.0facebook0.6B≈0.4 GB at 4-bit
- jina-embeddings-v3Jina AI0.6B≈0.4 GB at 4-bit
- xlm-roberta-largeFacebook AI community0.6B≈0.4 GB at 4-bit
- GOT-OCR-2.0-hfstepfun-ai0.6B≈0.4 GB at 4-bit
- bge-reranker-largeBAAI0.6B≈0.4 GB at 4-bit
- Multilingual-e5-large-instructintfloat0.6B≈0.4 GB at 4-bit4 also selling it hosted
- bloom-560mBigScience Workshop0.6B≈0.4 GB at 4-bit
- MiraTTSYatharthS0.5B≈0.3 GB at 4-bit
- Fun-CosyVoice3-0.5B-2512QwenAudio0.5B≈0.3 GB at 4-bit
- Qwen1.5-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- VoxCPM-0.5BOpenBMB0.5B≈0.3 GB at 4-bit
- reader-lm-0.5bJina AI0.5B≈0.3 GB at 4-bit
- Qwen2-0.5B-InstructAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2.5-0.5BAlibaba0.5B≈0.3 GB at 4-bit
- Qwen2.5-0.5B-InstructAlibaba0.5B≈0.3 GB at 4-bit1 also selling it hosted
- nomic-embed-text-v2-moeNomic AI0.5B≈0.3 GB at 4-bit
- LaBSEsentence-transformers0.5B≈0.3 GB at 4-bit
- blip-image-captioning-largeSalesforce0.5B≈0.3 GB at 4-bit
- LFM2.5-VL-450MLiquid AI0.4B≈0.3 GB at 4-bit
- PairRMLLM Blender0.4B≈0.3 GB at 4-bit
- stella_en_400M_v5NovaSearch0.4B≈0.3 GB at 4-bit
- gte-large-en-v1.5Alibaba-NLP0.4B≈0.3 GB at 4-bit
- CLIP-GmP-ViT-L-14zer0int0.4B≈0.3 GB at 4-bit
- circuit-sparsityOpenAI0.4B≈0.3 GB at 4-bit
- bart-large-cnnAI at Meta0.4B≈0.3 GB at 4-bit
- DolphinByteDance0.4B≈0.3 GB at 4-bit
- ModernBERT-largeAnswer.AI0.4B≈0.3 GB at 4-bit
- blip-vqa-baseSalesforce0.4B≈0.3 GB at 4-bit
- control_v1p_sd15_brightnessLatent Cat0.4B≈0.2 GB at 4-bit
- controlnet_qrcode-control_v1p_sd15DionTimmer0.4B≈0.2 GB at 4-bit
- LFM2-350MLiquid AI0.4B≈0.2 GB at 4-bit
- LFM2.5-350MLiquid AI0.4B≈0.2 GB at 4-bit
- IndicF5AI4Bharat0.4B≈0.2 GB at 4-bit
- nougat-basefacebook0.3B≈0.2 GB at 4-bit
- FLUX.1-Turbo-Alphaalimama-creative0.3B≈0.2 GB at 4-bit
- bert-large-uncased-whole-word-masking-finetuned-squadBERT community0.3B≈0.2 GB at 4-bit
- bge-large-enBAAI0.3B≈0.2 GB at 4-bit1 also selling it hosted
- UAE-Large-V1WhereIsAI0.3B≈0.2 GB at 4-bit
- mxbai-embed-large-v1Mixedbread0.3B≈0.2 GB at 4-bit
- trocr-base-handwrittenMicrosoft0.3B≈0.2 GB at 4-bit
- bge-large-zhBAAI0.3B≈0.2 GB at 4-bit
- text2vec-large-chineseGanymedeNil0.3B≈0.2 GB at 4-bit
- VoiceCraftpyp10.3B≈0.2 GB at 4-bit
- VibeVoice-ASR-BitNetMicrosoft0.3B≈0.2 GB at 4-bit
- wav2vec2-lg-xlsr-en-speech-emotion-recognitionehcalabres0.3B≈0.2 GB at 4-bit
- wav2vec2-large-xlsr-53-englishjonatasgrosman0.3B≈0.2 GB at 4-bit
- gte-multilingual-baseAlibaba-NLP0.3B≈0.2 GB at 4-bit
- dinov3-vitl16-pretrain-lvd1689mAI at Meta0.3B≈0.2 GB at 4-bit
- xlm-roberta-baseFacebook AI community0.3B≈0.2 GB at 4-bit
- mdeberta-v3-base-squad2timpal0l0.3B≈0.2 GB at 4-bit
- multilingual-e5-baseintfloat0.3B≈0.2 GB at 4-bit
- paraphrase-multilingual-mpnet-base-v2sentence-transformers0.3B≈0.2 GB at 4-bit
- functiongemma-270m-itGoogle0.3B≈0.2 GB at 4-bit
- gemma-3-270mGoogle0.3B≈0.2 GB at 4-bit
- gemma-3-270m-itGoogle0.3B≈0.2 GB at 4-bit
- granite-docling-258MIBM Granite0.3B≈0.2 GB at 4-bit
- NeoBERTChandar Research Lab0.2B≈0.2 GB at 4-bit
- whisper-smallOpenAI0.2B≈0.2 GB at 4-bit
- Florence-2-baseMicrosoft0.2B≈0.2 GB at 4-bit
- t5-baseT5 community0.2B≈0.1 GB at 4-bit
- chatgpt_paraphraser_on_T5_baseHumarin0.2B≈0.1 GB at 4-bit
- jina-clip-v1Jina AI0.2B≈0.1 GB at 4-bit
- BiRefNetZhengPeng70.2B≈0.1 GB at 4-bit
- RMBG-2.0BRIA AI0.2B≈0.1 GB at 4-bit
- bert-base-multilingual-casedBERT community0.2B≈0.1 GB at 4-bit
- Audio8-TTS-Preview-0.1bEdge00.2B≈0.1 GB at 4-bit
- jina-embeddings-v2-base-zhJina AI0.2B≈0.1 GB at 4-bit
- gpt-neo-125mEleutherAI0.2B≈0.1 GB at 4-bit
- ModernBERT-baseAnswer.AI0.1B≈0.1 GB at 4-bit
- gte-modernbert-baseAlibaba-NLP0.1B≈0.1 GB at 4-bit
- modernbert-embed-baseNomic AI0.1B≈0.1 GB at 4-bit
- bart-basefacebook0.1B≈0.1 GB at 4-bit
- jina-embeddings-v2-base-enJina AI0.1B≈0.1 GB at 4-bit
- MagicPrompt-Stable-DiffusionGustavosta0.1B≈0.1 GB at 4-bit
- gpt2OpenAI community0.1B≈0.1 GB at 4-bit
- nomic-embed-text-v1Nomic AI0.1B≈0.1 GB at 4-bit
- distilbert-base-multilingual-caseddistilbert0.1B≈0.1 GB at 4-bit
- distiluse-base-multilingual-cased-v2sentence-transformers0.1B≈0.1 GB at 4-bit
- SmolLM2-135MHuggingFaceTB0.1B≈0.1 GB at 4-bit
- roberta-baseFacebook AI community0.1B≈0.1 GB at 4-bit
- roberta-base-squad2deepset0.1B≈0.1 GB at 4-bit
- openjourneyprompthero0.1B≈0.1 GB at 4-bit
- paraphrase-multilingual-MiniLM-L12-v2sentence-transformers0.1B≈0.1 GB at 4-bit
- bert-base-uncasedBERT community0.1B≈0.1 GB at 4-bit
- pubmedbert-base-embeddingsNeuML0.1B≈0.1 GB at 4-bit
- bert-base-casedBERT community0.1B≈0.1 GB at 4-bit
- bert-base-chineseBERT community0.1B≈0.1 GB at 4-bit
- BEN2Prama LLC0.1B≈0.1 GB at 4-bit
- wav2vec2-base-960hAI at Meta0.1B≈0.1 GB at 4-bit
- nomic-embed-vision-v1.5Nomic AI0.1B≈0.1 GB at 4-bit
- distilgpt2DistilBERT community0.1B≈0.1 GB at 4-bit
- ast-finetuned-audioset-10-10-0.4593Massachusetts Institute of Technology0.1B≈0.1 GB at 4-bit
- dinov2-basefacebook0.1B≈0.1 GB at 4-bit
- vit-base-patch16-224Google0.1B≈0.1 GB at 4-bit
- vit-base-patch16-224-in21kGoogle0.1B≈0.1 GB at 4-bit