Conversation models that run on 128 GB
613 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 613 in all, a hundred to a page; this is page 3 of 7.
A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- Llama2-13B-TiefighterKoboldAI13.0B≈8.5 GB at 4-bit1 also selling it hosted
- Llama2-Chinese-13b-ChatFlagAlpha13.0B≈8.5 GB at 4-bit
- MythoMax 13BGryphe13.0B≈8.5 GB at 4-bit4 also selling it hosted
- Nous Hermes Llama2 13B13.0B≈8.5 GB at 4-bit2 also selling it hosted
- OpenOrca-Platypus2-13BOpenOrca13.0B≈8.5 GB at 4-bit
- Orca-2-13bMicrosoft13.0B≈8.5 GB at 4-bit
- ReMM SLERP 13BUndi9513.0B≈8.5 GB at 4-bit1 also selling it hosted
- WhiteRabbitNeo-13B-v1WhiteRabbitNeo13.0B≈8.5 GB at 4-bit
- Wizard-Vicuna-13B-UncensoredQuixi AI13.0B≈8.5 GB at 4-bit
- Wizard-Vicuna-13B-Uncensored-HFTheBloke13.0B≈8.5 GB at 4-bit
- WizardLM-13B-UncensoredQuixi AI13.0B≈8.5 GB at 4-bit
- WizardLM-13B-V1.2WizardLM Team13.0B≈8.5 GB at 4-bit
- Ziya-LLaMA-13B-v1Fengshenbang-LM13.0B≈8.5 GB at 4-bit
- chronos-hermes-13b-v2Austism13.0B≈8.5 GB at 4-bit2 also selling it hosted
- jais-13binception4213.0B≈8.5 GB at 4-bit
- jais-13b-chatinception4213.0B≈8.5 GB at 4-bit1 also selling it hosted
- llama-13bhuggyllama13.0B≈8.5 GB at 4-bit
- mythalion-13bPygmalionAI13.0B≈8.5 GB at 4-bit
- open_llama_13bOpenLM Research13.0B≈8.5 GB at 4-bit
- prometheus-13b-v1.0prometheus-eval13.0B≈8.5 GB at 4-bit
- ruGPT-3.5-13Bai-forever13.0B≈8.5 GB at 4-bit
- stable-vicuna-13b-deltaCarperAI13.0B≈8.5 GB at 4-bit
- vicuna-13b-v1.5Large Model Systems Organization13.0B≈8.5 GB at 4-bit
- vicuna-13b-v1.5-16kLarge Model Systems Organization13.0B≈8.5 GB at 4-bit
- Wayfarer-12BLatitude12.2B≈8.0 GB at 4-bit
- Mistral NemoMistral AI12.2B≈8.0 GB at 4-bit10 also selling it hosted
- MN-12B-Celeste-V1.9nothingiisreal12.0B≈7.8 GB at 4-bit
- NVIDIA-Nemotron-Nano-12B-v2NVIDIA12.0B≈7.8 GB at 4-bit2 also selling it hosted
- NemoMix-Unleashed-12BMarinaraSpaghetti12.0B≈7.8 GB at 4-bit
- Pixtral 12B 2409Mistral AI12.0B≈7.8 GB at 4-bit3 also selling it hosted
- gemma-3-12b-it-qat-q4_0-unquantizedGoogle12.0B≈7.8 GB at 4-bit
- oasst-sft-1-pythia-12bOpenAssistant12.0B≈7.8 GB at 4-bit
- oasst-sft-4-pythia-12b-epoch-3.5OpenAssistant12.0B≈7.8 GB at 4-bit
- pythia-12bEleutherAI12.0B≈7.8 GB at 4-bit1 also selling it hosted
- falcon-11BTechnology Innovation Institute11.1B≈7.2 GB at 4-bit
- Bielik-11B-v3.0-Instructspeakleash11.0B≈7.2 GB at 4-bit1 also selling it hosted
- Llama 3.2 11B Vision InstructMeta11.0B≈7.2 GB at 4-bit7 also selling it hosted
- Llama-3.2V-11B-cotXkev11.0B≈7.2 GB at 4-bit
- YanoljaNEXT-EEVE-Instruct-10.8Byanolja10.8B≈7.0 GB at 4-bit
- HyperCLOVAX-SEED-Omni-8BHyperCLOVA X10.7B≈7.0 GB at 4-bit
- Fimbulvetr-11B-v2Sao10K10.7B≈7.0 GB at 4-bit
- Solar-10.7B-Instruct-v1.0Upstage10.7B≈7.0 GB at 4-bit
- Solar-10.7B-v1.0Upstage10.7B≈7.0 GB at 4-bit
- Llama-3.2-11B-VisionMeta Llama10.6B≈6.9 GB at 4-bit
- Step3-VL-10BStepFun10.2B≈6.6 GB at 4-bit
- Qwen3.8-9B-Distillempero-ai9.7B≈6.3 GB at 4-bit
- fuyu-8bAdept AI Labs9.4B≈6.1 GB at 4-bit
- moondream3-previewmoondream9.3B≈6.0 GB at 4-bit
- gemma-2-9bGoogle9.2B≈6.0 GB at 4-bit
- EuroLLM-9B-InstructUTTER - Unified Transcription and Translation for Extended Reality9.2B≈5.9 GB at 4-bit
- EuroLLM-9BUTTER - Unified Transcription and Translation for Extended Reality9.0B≈5.9 GB at 4-bit
- Ornith-1.5-9Bornith-ai9.0B≈5.9 GB at 4-bit
- Ovis1.6-Gemma2-9BATH-MaaS9.0B≈5.9 GB at 4-bit
- Ovis2.5-9BATH-MaaS9.0B≈5.9 GB at 4-bit
- Qwen3.5-9B-BaseAlibaba9.0B≈5.9 GB at 4-bit1 also selling it hosted
- Qwythos-9B-v2empero-ai9.0B≈5.9 GB at 4-bit
- Turkish-Gemma-9b-T1ytu-ce-cosmos9.0B≈5.9 GB at 4-bit
- Yi-1.5-9B-Chat01-ai9.0B≈5.9 GB at 4-bit
- gemma-2-9b-it-SimPOprinceton-nlp9.0B≈5.9 GB at 4-bit
- Carnice-9bkai-os9.0B≈5.8 GB at 4-bit
- NeoHorse-1-9BTokenRhythm9.0B≈5.8 GB at 4-bit
- Yi-9B01-ai8.8B≈5.7 GB at 4-bit
- internlm3-8b-instructinternlm8.8B≈5.7 GB at 4-bit
- Granite 4.1 8BIBM watsonx.ai8.8B≈5.7 GB at 4-bit
- Cosmos-Reason2-8BNVIDIA8.8B≈5.7 GB at 4-bit
- Huihui-Qwen3-VL-8B-Instruct-abliteratedhuihui-ai8.8B≈5.7 GB at 4-bit
- MAI-UI-8BTongyi-MAI8.8B≈5.7 GB at 4-bit
- Qwen3 VL 8B InstructAlibaba8.8B≈5.7 GB at 4-bit8 also selling it hosted
- gemma-7bGoogle8.5B≈5.5 GB at 4-bit1 also selling it hosted
- gemma-7b-itGoogle8.5B≈5.5 GB at 4-bit3 also selling it hosted
- jetmoe-8bJetMoE8.5B≈5.5 GB at 4-bit
- LFM2.5-8B-A1BLiquid AI8.5B≈5.5 GB at 4-bit
- idefics2-8bHuggingFaceM48.4B≈5.5 GB at 4-bit
- LFM2-8B-A1BLiquid AI8.3B≈5.4 GB at 4-bit
- Cosmos-Reason1-7BNVIDIA8.3B≈5.4 GB at 4-bit
- Fara-7BMicrosoft8.3B≈5.4 GB at 4-bit
- Holo1-7BH Company8.3B≈5.4 GB at 4-bit
- Qwen2.5-VL-7B-InstructAlibaba8.3B≈5.4 GB at 4-bit1 also selling it hosted
- Qwen2-VL-7B-InstructAlibaba8.3B≈5.4 GB at 4-bit1 also selling it hosted
- UI-TARS-7B-DPOByteDance Seed8.3B≈5.4 GB at 4-bit
- zeta-2Zed Industries8.3B≈5.4 GB at 4-bit
- VLM_WebSight_finetunedHuggingFaceM48.2B≈5.3 GB at 4-bit
- II-Medical-8BIntelligent Internet8.2B≈5.3 GB at 4-bit
- MiniCPM4.1-8BOpenBMB8.2B≈5.3 GB at 4-bit
- granite-3.0-8b-instructIBM Granite8.2B≈5.3 GB at 4-bit1 also selling it hosted
- dolphin-2.9-llama3-8bDolphin8.0B≈5.2 GB at 4-bit
- UserLM-8bMicrosoft8.0B≈5.2 GB at 4-bit
- Llama3-ChatQA-1.5-8BNVIDIA8.0B≈5.2 GB at 4-bit
- DarkIdol-Llama-3.1-8B-Instruct-1.2-Uncensoredaifeifei7988.0B≈5.2 GB at 4-bit
- KernelLLMfacebook8.0B≈5.2 GB at 4-bit
- Llama-3-8B-Instruct-Gradient-1048kDeepSky8.0B≈5.2 GB at 4-bit
- Llama-3-8B-WebMcGill NLP Group8.0B≈5.2 GB at 4-bit
- Llama-3.1-Nemotron-Nano-8B-v1NVIDIA8.0B≈5.2 GB at 4-bit
- Llama-3.1-SuperNova-LiteArcee AI8.0B≈5.2 GB at 4-bit
- Llama3-8B-Chinese-Chatshenzhi-wang8.0B≈5.2 GB at 4-bit
- Meta-Llama-3.1-8B-Instruct-abliteratedmlabonne8.0B≈5.2 GB at 4-bit
- llama-3-Korean-Bllossom-8BMLP-LAB8.0B≈5.2 GB at 4-bit
- aya-23-8BCohere Labs8.0B≈5.2 GB at 4-bit
- c4ai-command-r7b-12-2024Cohere Labs8.0B≈5.2 GB at 4-bit
- Molmo-7B-D-0924Ai28.0B≈5.2 GB at 4-bit