Models that run on 128 GB
1133 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 1,133 in all, a hundred to a page; this is page 5 of 12.
A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.
- Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSOREDDavidAU9.4B≈6.1 GB at 4-bit
- Qwythos-9B-Claude-Mythos-5-1Mempero-ai9.4B≈6.1 GB at 4-bit
- fuyu-8bAdept AI Labs9.4B≈6.1 GB at 4-bit
- MiniCPM-o-4_5OpenBMB9.4B≈6.1 GB at 4-bit
- VibeVoice-7BVibeVoice Community (Unofficial)9.3B≈6.1 GB at 4-bit
- VibeVoice-Largeaoi-ot9.3B≈6.1 GB at 4-bit
- kugelaudio-0-openKugelaudio9.3B≈6.1 GB at 4-bit
- ideogram-4-fp8Ideogram9.3B≈6.0 GB at 4-bit
- moondream3-previewmoondream9.3B≈6.0 GB at 4-bit
- bge-multilingual-gemma2BAAI9.2B≈6.0 GB at 4-bit
- gemma-2-9bGoogle9.2B≈6.0 GB at 4-bit
- EuroLLM-9B-InstructUTTER - Unified Transcription and Translation for Extended Reality9.2B≈5.9 GB at 4-bit
- FLUX.2-klein-9b-kvBlack Forest Labs9.1B≈5.9 GB at 4-bit
- AutoGLM-Phone-9B-MultilingualZ.ai9.0B≈5.9 GB at 4-bit3 also selling it hosted
- EuroLLM-9BUTTER - Unified Transcription and Translation for Extended Reality9.0B≈5.9 GB at 4-bit
- FLUX.2-klein-9b-fp8Black Forest Labs9.0B≈5.9 GB at 4-bit
- FLUX.2-klein-base-9BBlack Forest Labs9.0B≈5.9 GB at 4-bit
- FLUX.2-klein-base-9b-fp8Black Forest Labs9.0B≈5.9 GB at 4-bit
- Flux2-Klein-9B-Enhanced-Detailsdx81529.0B≈5.9 GB at 4-bit
- Flux2-Klein-9B-True-V2wikeeyang9.0B≈5.9 GB at 4-bit
- Flux2-Klein-9B-True-V3wikeeyang9.0B≈5.9 GB at 4-bit
- Nemotron Nano 9B V2NVIDIA9.0B≈5.9 GB at 4-bit5 also selling it hosted
- Ornith-1.5-9Bornith-ai9.0B≈5.9 GB at 4-bit
- Ovis1.6-Gemma2-9BATH-MaaS9.0B≈5.9 GB at 4-bit
- Ovis2.5-9BATH-MaaS9.0B≈5.9 GB at 4-bit
- Qwen3.5-9B-BaseAlibaba9.0B≈5.9 GB at 4-bit1 also selling it hosted
- Qwythos-9B-v2empero-ai9.0B≈5.9 GB at 4-bit
- Turkish-Gemma-9b-T1ytu-ce-cosmos9.0B≈5.9 GB at 4-bit
- Yi-1.5-9B-Chat01-ai9.0B≈5.9 GB at 4-bit
- codegeex4-all-9bzai-org9.0B≈5.9 GB at 4-bit
- gemma-2-9b-itGoogle9.0B≈5.9 GB at 4-bit3 also selling it hosted
- gemma-2-9b-it-SimPOprinceton-nlp9.0B≈5.9 GB at 4-bit
- Carnice-9bkai-os9.0B≈5.8 GB at 4-bit
- NeoHorse-1-9BTokenRhythm9.0B≈5.8 GB at 4-bit
- Chroma1-HDlodestones8.9B≈5.8 GB at 4-bit
- Yi-9B01-ai8.8B≈5.7 GB at 4-bit
- Yi-Coder-9B-Chat01-ai8.8B≈5.7 GB at 4-bit
- internlm3-8b-instructinternlm8.8B≈5.7 GB at 4-bit
- Granite 4.1 8BIBM watsonx.ai8.8B≈5.7 GB at 4-bit
- Cosmos-Reason2-8BNVIDIA8.8B≈5.7 GB at 4-bit
- Huihui-Qwen3-VL-8B-Instruct-abliteratedhuihui-ai8.8B≈5.7 GB at 4-bit
- MAI-UI-8BTongyi-MAI8.8B≈5.7 GB at 4-bit
- Qwen3 VL 8B InstructAlibaba8.8B≈5.7 GB at 4-bit8 also selling it hosted
- Qwen3 VL 8B ThinkingAlibaba8.8B≈5.7 GB at 4-bit3 also selling it hosted
- chandraDatalab8.8B≈5.7 GB at 4-bit
- MiniCPM-o-2_6OpenBMB8.7B≈5.6 GB at 4-bit
- VibeVoice-ASRMicrosoft8.7B≈5.6 GB at 4-bit
- VibeVoice-ASR-Streaming-7BMicrosoft8.7B≈5.6 GB at 4-bit
- Molmo2-8BAllen Institute for AI (Ai2)8.7B≈5.6 GB at 4-bit
- codegemma-7bGoogle8.5B≈5.5 GB at 4-bit1 also selling it hosted
- gemma-7bGoogle8.5B≈5.5 GB at 4-bit1 also selling it hosted
- gemma-7b-itGoogle8.5B≈5.5 GB at 4-bit3 also selling it hosted
- jetmoe-8bJetMoE8.5B≈5.5 GB at 4-bit
- Emu3-GenBAAI8.5B≈5.5 GB at 4-bit
- MOSS-TTSOpenMOSS-Team8.5B≈5.5 GB at 4-bit
- MOSS-TTS-v1.5OpenMOSS-Team8.5B≈5.5 GB at 4-bit
- LFM2.5-8B-A1BLiquid AI8.5B≈5.5 GB at 4-bit
- idefics2-8bHuggingFaceM48.4B≈5.5 GB at 4-bit
- personaplex-7b-v1NVIDIA8.4B≈5.4 GB at 4-bit
- LFM2-8B-A1BLiquid AI8.3B≈5.4 GB at 4-bit
- Cosmos-Reason1-7BNVIDIA8.3B≈5.4 GB at 4-bit
- Fara-7BMicrosoft8.3B≈5.4 GB at 4-bit
- Holo1-7BH Company8.3B≈5.4 GB at 4-bit
- NuMarkdown-8B-ThinkingNuMind8.3B≈5.4 GB at 4-bit
- Qwen2.5-VL-7B-InstructAlibaba8.3B≈5.4 GB at 4-bit1 also selling it hosted
- RolmOCRReducto8.3B≈5.4 GB at 4-bit1 also selling it hosted
- Qwen2-VL-7B-InstructAlibaba8.3B≈5.4 GB at 4-bit1 also selling it hosted
- UI-TARS-7B-DPOByteDance Seed8.3B≈5.4 GB at 4-bit
- olmOCR-7B-0225-previewAi28.3B≈5.4 GB at 4-bit
- zeta-2Zed Industries8.3B≈5.4 GB at 4-bit
- VLM_WebSight_finetunedHuggingFaceM48.2B≈5.3 GB at 4-bit
- II-Medical-8BIntelligent Internet8.2B≈5.3 GB at 4-bit
- Qwen3 8BAlibaba8.2B≈5.3 GB at 4-bit8 also selling it hosted
- MisoTTSMiso Labs8.2B≈5.3 GB at 4-bit
- MiniCPM4.1-8BOpenBMB8.2B≈5.3 GB at 4-bit
- granite-3.0-8b-instructIBM Granite8.2B≈5.3 GB at 4-bit1 also selling it hosted
- Flex.2-previewostris8.2B≈5.3 GB at 4-bit
- Flex.1-alphaostris8.2B≈5.3 GB at 4-bit
- flux.1-lite-8B-alphaFreepik8.2B≈5.3 GB at 4-bit
- Qwen3-VL-Embedding-8BAlibaba8.1B≈5.3 GB at 4-bit
- ERNIE-ImageBaidu8.0B≈5.2 GB at 4-bit
- ERNIE-Image-TurboBaidu8.0B≈5.2 GB at 4-bit
- dolphin-2.9-llama3-8bDolphin8.0B≈5.2 GB at 4-bit
- UserLM-8bMicrosoft8.0B≈5.2 GB at 4-bit
- Llama3-ChatQA-1.5-8BNVIDIA8.0B≈5.2 GB at 4-bit
- Aion-RP 1.0 (8B)AionLabs8.0B≈5.2 GB at 4-bit2 also selling it hosted
- DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensoredaifeifei7988.0B≈5.2 GB at 4-bit
- KernelLLMfacebook8.0B≈5.2 GB at 4-bit
- Llama 3.1 8B InstructMeta8.0B≈5.2 GB at 4-bit20 also selling it hosted
- Llama-3-8B-Instruct-Gradient-1048kDeepSky8.0B≈5.2 GB at 4-bit
- Llama-3-8B-WebMcGill NLP Group8.0B≈5.2 GB at 4-bit
- Llama-3-RefueledRefuel AI8.0B≈5.2 GB at 4-bit
- Llama-3.1-Nemotron-Nano-8B-v1NVIDIA8.0B≈5.2 GB at 4-bit
- Llama-3.1-SuperNova-LiteArcee AI8.0B≈5.2 GB at 4-bit
- Llama3-8B-Chinese-Chatshenzhi-wang8.0B≈5.2 GB at 4-bit
- Meta-Llama-3.1-8B-Instruct-abliteratedmlabonne8.0B≈5.2 GB at 4-bit
- llama-3-Korean-Bllossom-8BMLP-LAB8.0B≈5.2 GB at 4-bit
- aya-23-8BCohere Labs8.0B≈5.2 GB at 4-bit
- c4ai-command-r7b-12-2024Cohere Labs8.0B≈5.2 GB at 4-bit
- Molmo-7B-D-0924Ai28.0B≈5.2 GB at 4-bit