Conversation models that run on 128 GB
613 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 613 in all, a hundred to a page; this is page 2 of 7.
A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- TinyR1-32B-Previewqihoo36032.8B≈21.3 GB at 4-bit
- DeepSWE-PreviewAgentica32.8B≈21.3 GB at 4-bit
- KAT-DevKwaipilot32.8B≈21.3 GB at 4-bit
- GLM-Z1-32B-0414zai-org32.6B≈21.2 GB at 4-bit
- sarvam-30bSarvam AI32.2B≈20.9 GB at 4-bit
- Baichuan M2 32Bbaichuan32.0B≈20.8 GB at 4-bit2 also selling it hosted
- EXAONE-Deep-32BLGAI-EXAONE32.0B≈20.8 GB at 4-bit
- GLM-4-32B-0414Z.ai32.0B≈20.8 GB at 4-bit3 also selling it hosted
- Olmo-2-0325-32B-InstructAllen Institute for AI (Ai2)32.0B≈20.8 GB at 4-bit
- Olmo-3.1-32B-InstructAllen Institute for AI (Ai2)32.0B≈20.8 GB at 4-bit1 also selling it hosted
- OpenThinker-32Bopen-thoughts32.0B≈20.8 GB at 4-bit
- QwQ-R1984-32BVIDraft32.0B≈20.8 GB at 4-bit
- Qwen-SEA-LION-v4-32B-ITaisingapore32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen2.5-32BAlibaba32.0B≈20.8 GB at 4-bit1 also selling it hosted
- Qwen2.5-VL-32B-InstructAlibaba32.0B≈20.8 GB at 4-bit3 also selling it hosted
- QwenLong-L1-32BTongyi-Zhiwen32.0B≈20.8 GB at 4-bit
- Sky-T1-32B-PreviewNovaSky-AI32.0B≈20.8 GB at 4-bit1 also selling it hosted
- UIGEN-X-32B-0727Tesslate32.0B≈20.8 GB at 4-bit
- s1-32Bsimplescaling32.0B≈20.8 GB at 4-bit
- NVIDIA Nemotron 3.5 Lightning 30B A3BNVIDIA31.6B≈20.5 GB at 4-bit2 also selling it hosted
- Nemotron 3 Nano 30B A3BNVIDIA31.6B≈20.5 GB at 4-bit10 also selling it hosted
- Nemotron-Cascade-2-30B-A3BNVIDIA31.6B≈20.5 GB at 4-bit
- phonellm-alpha-1Pipecat31.6B≈20.5 GB at 4-bit
- Qwen3 VL 30B A3B InstructAlibaba31.1B≈20.2 GB at 4-bit8 also selling it hosted
- Gemma-4-31B-it-pearlpearl-ai31.0B≈20.2 GB at 4-bit
- MiroThinker-v1.5-30BMiroMind AI30.5B≈19.8 GB at 4-bit
- Tongyi-DeepResearch-30B-A3BAlibaba-NLP30.5B≈19.8 GB at 4-bit
- Nemotron-Labs-Audex-30B-A3BNVIDIA30.0B≈19.5 GB at 4-bit
- Ovis2.6-30B-A3BATH-MaaS30.0B≈19.5 GB at 4-bit
- QwenLong-L1.5-30B-A3BTongyi-Zhiwen30.0B≈19.5 GB at 4-bit
- Wizard-Vicuna-30B-UncensoredQuixi AI30.0B≈19.5 GB at 4-bit
- WizardLM-30B-UncensoredQuixi AI30.0B≈19.5 GB at 4-bit
- granite-4.1-30bIBM Granite30.0B≈19.5 GB at 4-bit
- Muse Glimmer 30BMeta29.8B≈19.4 GB at 4-bit8 also selling it hosted
- medgemma-27b-itGoogle28.8B≈18.7 GB at 4-bit
- ERNIE 4.5 VL 28B A3BBaidu28.0B≈18.2 GB at 4-bit2 also selling it hosted
- Huihui-Qwen3.8-27B-abliteratedhuihui-ai27.8B≈18.1 GB at 4-bit
- Qwen3.8 27BAlibaba27.8B≈18.1 GB at 4-bit12 also selling it hosted
- Qwen3.8-27B-AEON-ULTIMATE-UNCENSORED-BF16AEON-727.8B≈18.1 GB at 4-bit
- Qwen3.8-27B-TURBO-Fable-Cold-Fusion-735-882-Heretic-Uncensored-NM-DAUDavidAU27.8B≈18.1 GB at 4-bit
- Qwen3.8-27B-UncensoredOrcaRouter27.8B≈18.1 GB at 4-bit
- Swift-Qwen3.8-27bukisai27.8B≈18.1 GB at 4-bit
- deepseek-vl2DeepSeek27.5B≈17.9 GB at 4-bit
- Gemma-3-R1984-27BVIDraft27.4B≈17.8 GB at 4-bit
- gemma-3-27b-it-abliteratedmlabonne27.4B≈17.8 GB at 4-bit
- datagemma-rag-27b-itGoogle27.2B≈17.7 GB at 4-bit
- medgemma-27b-text-itGoogle27.0B≈17.6 GB at 4-bit
- C2S-Scale-Gemma-2-27Bvandijklab27.0B≈17.6 GB at 4-bit
- Fara1.5-27BMicrosoft27.0B≈17.6 GB at 4-bit
- Gemma-SEA-LION-v4-27B-ITaisingapore27.0B≈17.6 GB at 4-bit2 also selling it hosted
- Huihui-Qwen3.5-27B-abliteratedhuihui-ai27.0B≈17.6 GB at 4-bit
- Qwen3.6-27B-AEON-Ultimate-Uncensored-BF16AEON-727.0B≈17.6 GB at 4-bit
- Qwen3.8-27B-DFlash2Z Lab27.0B≈17.6 GB at 4-bit
- Qwythos-27B-v1empero-ai27.0B≈17.6 GB at 4-bit
- Gemma 4 26B A4B ITGoogle26.0B≈16.9 GB at 4-bit
- diffusiongemma-26B-A4B-itGoogle25.8B≈16.8 GB at 4-bit
- Gemma 4 26B A4BGoogle25.8B≈16.8 GB at 4-bit7 also selling it hosted
- AriaRhymes.AI25.3B≈16.4 GB at 4-bit
- Dolphin Mistral 24B Venice EditionCognitive Computations24.0B≈15.6 GB at 4-bit1 also selling it hosted
- Mistral Small 3.1 24B Instruct 2503Mistral AI24.0B≈15.6 GB at 4-bit3 also selling it hosted
- Mistral-Small-3.2-24B-Instruct-2506Mistral AI24.0B≈15.6 GB at 4-bit1 also selling it hosted
- Mistral Small 3Mistral AI23.6B≈15.3 GB at 4-bit4 also selling it hosted
- sarvam-mSarvam AI23.6B≈15.3 GB at 4-bit
- solar-pro-preview-instructUpstage22.1B≈14.4 GB at 4-bit
- gpt-oss-safeguard-20bOpenAI21.5B≈14.0 GB at 4-bit6 also selling it hosted
- ERNIE-4.5-21B-A3B-PTBaidu21.0B≈13.7 GB at 4-bit1 also selling it hosted
- context-1chroma20.9B≈13.6 GB at 4-bit
- gpt-neox-20bEleutherAI20.7B≈13.5 GB at 4-bit
- maple-previewdeepgrove20.2B≈13.1 GB at 4-bit
- Gemma-4-31B-JANG_4M-CRACKdealignai20.2B≈13.1 GB at 4-bit
- cogvlm2-llama3-chat-19Bzai-org19.5B≈12.7 GB at 4-bit
- cogvlm-chat-hfzai-org17.6B≈11.5 GB at 4-bit
- Ling-mini-2.0InclusionAI16.3B≈10.6 GB at 4-bit
- Ring-mini-2.0InclusionAI16.3B≈10.6 GB at 4-bit
- Instella-MoE-16B-A3B-Thinkamd16.0B≈10.4 GB at 4-bit
- deepseek-moe-16b-basedeepseek-ai16.0B≈10.4 GB at 4-bit
- deepseek-moe-16b-chatdeepseek-ai16.0B≈10.4 GB at 4-bit
- Moonlight-16B-A3B-InstructMoonshot AI16.0B≈10.4 GB at 4-bit
- DeepSeek-V2-Litedeepseek-ai15.7B≈10.2 GB at 4-bit
- starchat-alphaHugging Face H415.5B≈10.1 GB at 4-bit
- Apriel-1.6-15b-ThinkerServiceNow-AI15.0B≈9.8 GB at 4-bit
- Apriel-1.5-15b-ThinkerServiceNow-AI14.9B≈9.7 GB at 4-bit
- Qwen2.5-14B-Instruct-1MAlibaba14.8B≈9.6 GB at 4-bit
- SuperNova-MediusArcee AI14.8B≈9.6 GB at 4-bit
- Qwen1.5-MoE-A2.7BAlibaba14.3B≈9.3 GB at 4-bit
- 14BCausalLM14.0B≈9.1 GB at 4-bit
- ChatTS-14Bbytedance-research14.0B≈9.1 GB at 4-bit
- Fathom-R1-14BFractalAIResearch14.0B≈9.1 GB at 4-bit
- Nemotron-Labs-Diffusion-14BNVIDIA14.0B≈9.1 GB at 4-bit
- Qwen2.5-14BAlibaba14.0B≈9.1 GB at 4-bit1 also selling it hosted
- Velvet-14BAlmawave14.0B≈9.1 GB at 4-bit
- rwkv-4-pile-14bBlinkDL14.0B≈9.1 GB at 4-bit
- miniGCausalLM14.0B≈9.1 GB at 4-bit
- NexusRaven-V2-13BNexusflow13.0B≈8.5 GB at 4-bit
- Llama-2-13b-hfMeta Llama13.0B≈8.5 GB at 4-bit
- Baichuan2-13B-ChatBaichuan Intelligent Technology13.0B≈8.5 GB at 4-bit
- LLaVA-13b-delta-v0liuhaotian13.0B≈8.5 GB at 4-bit
- Llama-2-13bMeta Llama13.0B≈8.5 GB at 4-bit
- Llama-2-13b-chatMeta Llama13.0B≈8.5 GB at 4-bit1 also selling it hosted
- Llama-2-13b-chat-hfMeta13.0B≈8.5 GB at 4-bit1 also selling it hosted