Conversation models that run on 128 GB
613 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 613 in all, a hundred to a page; this is page 4 of 7.
A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- LLaDA-8B-InstructGSAI-ML8.0B≈5.2 GB at 4-bit
- Apertus-8B-2509swiss-ai8.0B≈5.2 GB at 4-bit
- Apertus-8B-Instruct-2509swiss-ai8.0B≈5.2 GB at 4-bit1 also selling it hosted
- Apertus-v1.5-8Bswiss-ai8.0B≈5.2 GB at 4-bit1 also selling it hosted
- Bio-Medical-MultiModal-Llama-3-8B-V1ContactDoctor8.0B≈5.2 GB at 4-bit
- Foundation-Sec-8Bfdtn-ai8.0B≈5.2 GB at 4-bit
- Hermes 2 Pro Llama 3 8BNous Research8.0B≈5.2 GB at 4-bit2 also selling it hosted
- Idefics3-8B-Llama3HuggingFaceM48.0B≈5.2 GB at 4-bit
- L3 8B Stheno V3.2Sao10K8.0B≈5.2 GB at 4-bit4 also selling it hosted
- L3-8B-Lunaris-v1Sao10K8.0B≈5.2 GB at 4-bit2 also selling it hosted
- Llama-3-8B-Lexi-UncensoredOrenguteng8.0B≈5.2 GB at 4-bit
- Llama-3-ELYZA-JP-8Belyza8.0B≈5.2 GB at 4-bit
- Llama-3-Groq-8B-Tool-UseGroq8.0B≈5.2 GB at 4-bit
- Llama-3-Open-Ko-8Bbeomi8.0B≈5.2 GB at 4-bit
- Llama-3.1-8B-Lexi-Uncensored-V2Orenguteng8.0B≈5.2 GB at 4-bit1 also selling it hosted
- Llama-3.1-Nemotron-Nano-VL-8B-V1NVIDIA8.0B≈5.2 GB at 4-bit
- Llama-3.1-Storm-8Bakjindal532448.0B≈5.2 GB at 4-bit
- Llama-3.1-Tulu-3-8BAllen Institute for AI (Ai2)8.0B≈5.2 GB at 4-bit
- Llama3-OpenBioLLM-8Baaditya8.0B≈5.2 GB at 4-bit
- Llama3-TAIDE-LX-8B-Chat-Alpha1taide8.0B≈5.2 GB at 4-bit
- Meta-Llama-Guard-2-8BMeta Llama8.0B≈5.2 GB at 4-bit1 also selling it hosted
- MiniCPM4-8BOpenBMB8.0B≈5.2 GB at 4-bit
- Mistral-NeMo-Minitron-8B-BaseNVIDIA8.0B≈5.2 GB at 4-bit
- NeuralDaredevil-8B-abliteratedmlabonne8.0B≈5.2 GB at 4-bit
- WeDLM-8B-InstructTencent Hunyuan8.0B≈5.2 GB at 4-bit
- aya-vision-8bCohere Labs8.0B≈5.2 GB at 4-bit
- deepthought-8b-llama-v0.01-alpharuliad8.0B≈5.2 GB at 4-bit
- granite-3.1-8b-instructIBM Granite8.0B≈5.2 GB at 4-bit
- granite-3.3-8b-instructIBM Granite8.0B≈5.2 GB at 4-bit1 also selling it hosted
- Ling-3.0-tinyInclusionAI7.9B≈5.1 GB at 4-bit
- EXAONE-3.0-7.8B-InstructLG AI Research7.8B≈5.1 GB at 4-bit
- phixtral-4x2_8mlabonne7.8B≈5.1 GB at 4-bit
- EXAONE-3.5-7.8B-InstructLGAI-EXAONE7.8B≈5.1 GB at 4-bit
- internlm2_5-7b-chatinternlm7.7B≈5.0 GB at 4-bit
- Qwen-7B-ChatAlibaba7.7B≈5.0 GB at 4-bit
- Qwen1.5-7B-ChatAlibaba7.7B≈5.0 GB at 4-bit
- DeepHat-V1-7BDeepHat7.6B≈5.0 GB at 4-bit
- Marco-o1ATH-MaaS7.6B≈5.0 GB at 4-bit
- Qwen2.5-7B-Instruct-1MAlibaba7.6B≈5.0 GB at 4-bit
- VulnLLM-R-7BVirtueAI7.6B≈5.0 GB at 4-bit
- falcon-H1R-7BTechnology Innovation Institute7.6B≈4.9 GB at 4-bit
- llava-v1.6-mistral-7bliuhaotian7.6B≈4.9 GB at 4-bit
- starvector-8b-im2svgstarvector7.5B≈4.9 GB at 4-bit
- falcon-mamba-7bTechnology Innovation Institute7.3B≈4.7 GB at 4-bit
- Starling-LM-7B-alphaBerkeley-Nest7.2B≈4.7 GB at 4-bit
- Starling-LM-7B-betaNexusflow7.2B≈4.7 GB at 4-bit
- dolphin-2.1-mistral-7bDolphin7.2B≈4.7 GB at 4-bit
- dolphin-2.2.1-mistral-7bDolphin7.2B≈4.7 GB at 4-bit
- dolphin-2.8-mistral-7b-v02Dolphin7.2B≈4.7 GB at 4-bit
- openchat-3.5-0106openchat7.2B≈4.7 GB at 4-bit
- Mistral-7B-v0.1Mistral AI7.2B≈4.7 GB at 4-bit
- neural-chat-7b-v3-1Intel7.2B≈4.7 GB at 4-bit
- zephyr-7b-alphaHugging Face H47.2B≈4.7 GB at 4-bit
- zephyr-7b-betaHugging Face H47.2B≈4.7 GB at 4-bit1 also selling it hosted
- falcon-7bTechnology Innovation Institute7.2B≈4.7 GB at 4-bit
- falcon-7b-instructTechnology Innovation Institute7.2B≈4.7 GB at 4-bit
- bloom-7b1BigScience Workshop7.1B≈4.6 GB at 4-bit
- llava-1.5-7b-hfLlava Hugging Face7.1B≈4.6 GB at 4-bit
- DeciLM-7BDeci AI7.0B≈4.6 GB at 4-bit
- chameleon-7bfacebook7.0B≈4.6 GB at 4-bit
- ALLaM-7B-Instruct-previewHUMAIN7.0B≈4.6 GB at 4-bit
- AlphaMonarch-7Bmlabonne7.0B≈4.5 GB at 4-bit
- BELLE-7B-2MBelleGroup7.0B≈4.5 GB at 4-bit
- BioMistral-7BBioMistral7.0B≈4.5 GB at 4-bit
- Chinese-Llama-2-7bLinkSoul7.0B≈4.5 GB at 4-bit
- DeepSeek-R1-Distill-QWEN-7BDeepSeek7.0B≈4.5 GB at 4-bit3 also selling it hosted
- Dream-v0-Instruct-7BDream-org7.0B≈4.5 GB at 4-bit
- FastVLM-7BApple7.0B≈4.5 GB at 4-bit
- K2-Horizon-7BInstitute of Foundation Models7.0B≈4.5 GB at 4-bit
- Lily-Cybersecurity-7B-v0.2segolilylabs7.0B≈4.5 GB at 4-bit
- Llama-2-7bMeta Llama7.0B≈4.5 GB at 4-bit
- Llama-2-7b-chatMeta Llama7.0B≈4.5 GB at 4-bit
- Llama2-Chinese-7b-ChatFlagAlpha7.0B≈4.5 GB at 4-bit
- MiMo-7B-RLXiaomiMiMo7.0B≈4.5 GB at 4-bit
- MiMo-VL-7B-RLXiaomiMiMo7.0B≈4.5 GB at 4-bit
- Mistral-7B-Instruct-v0.1Mistral AI7.0B≈4.5 GB at 4-bit3 also selling it hosted
- Mistral-7B-Instruct-v0.2Mistral AI7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Mistral-7B-OpenOrcaOpenOrca7.0B≈4.5 GB at 4-bit
- Mistral-Trismegistus-7Bteknium7.0B≈4.5 GB at 4-bit
- Molmo-7B-O-0924Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit
- NeuralBeagle14-7Bmlabonne7.0B≈4.5 GB at 4-bit
- NeuralHermes-2.5-Mistral-7Bmlabonne7.0B≈4.5 GB at 4-bit
- OLMoE-1B-7B-0924Allen Institute for AI (Ai2)7.0B≈4.5 GB at 4-bit
- OpenHermes-2-Mistral-7Bteknium7.0B≈4.5 GB at 4-bit1 also selling it hosted
- Openhermes2.5 Mistral 7Bteknium7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Orca-2-7bMicrosoft7.0B≈4.5 GB at 4-bit
- Qwen 2 7B InstructAlibaba7.0B≈4.5 GB at 4-bit4 also selling it hosted
- Qwen2-7BAlibaba7.0B≈4.5 GB at 4-bit
- TAIDE-LX-7B-Chattaide7.0B≈4.5 GB at 4-bit
- UI-TARS-7B-SFTByteDance Seed7.0B≈4.5 GB at 4-bit
- Vistral-7B-ChatViet-Mistral7.0B≈4.5 GB at 4-bit
- WizardLM-7B-UncensoredQuixi AI7.0B≈4.5 GB at 4-bit
- chinese-alpaca-2-7bJoint Laboratory of HIT and iFLYTEK Research (HFL)7.0B≈4.5 GB at 4-bit
- deepseek-llm-7b-basedeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-llm-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-math-7b-instructdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-vl-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- gemma-1.1-7b-itGoogle7.0B≈4.5 GB at 4-bit1 also selling it hosted
- internlm-xcomposer2d5-7binternlm7.0B≈4.5 GB at 4-bit
- mixtral-7b-8expertDiscoResearch7.0B≈4.5 GB at 4-bit