Conversation models that run on 256 GB
628 models with published weights that fit in 256 GB — a Mac Studio with 256. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 274 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute. 628 in all, a hundred to a page; this is page 5 of 7.
A Mac Studio at the top of its configuration, or a serious machine. Almost everything openly published fits, including the large mixtures of experts. At this point the constraint is no longer whether the model loads but whether it generates fast enough to be worth waiting for.
Models that answer in prose across whatever subject you put to them. This is the largest category and the least differentiated: most of them are competent at most things, so the choice usually comes down to price, context length and how the maker behaves about availability. Look at what a model costs per million tokens in and out — the output rate is often four or five times the input rate, and it is the one that decides your bill.
- Orca-2-7bMicrosoft7.0B≈4.5 GB at 4-bit
- Qwen 2 7B InstructAlibaba7.0B≈4.5 GB at 4-bit4 also selling it hosted
- Qwen2-7BAlibaba7.0B≈4.5 GB at 4-bit
- TAIDE-LX-7B-Chattaide7.0B≈4.5 GB at 4-bit
- UI-TARS-7B-SFTByteDance Seed7.0B≈4.5 GB at 4-bit
- Vistral-7B-ChatViet-Mistral7.0B≈4.5 GB at 4-bit
- WizardLM-7B-UncensoredQuixi AI7.0B≈4.5 GB at 4-bit
- chinese-alpaca-2-7bJoint Laboratory of HIT and iFLYTEK Research (HFL)7.0B≈4.5 GB at 4-bit
- deepseek-llm-7b-basedeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-llm-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-math-7b-instructdeepseek-ai7.0B≈4.5 GB at 4-bit
- deepseek-vl-7b-chatdeepseek-ai7.0B≈4.5 GB at 4-bit
- gemma-1.1-7b-itGoogle7.0B≈4.5 GB at 4-bit1 also selling it hosted
- internlm-xcomposer2d5-7binternlm7.0B≈4.5 GB at 4-bit
- mixtral-7b-8expertDiscoResearch7.0B≈4.5 GB at 4-bit
- open-calm-7bCyberAgent7.0B≈4.5 GB at 4-bit
- rwkv-4-pile-7bBlinkDL7.0B≈4.5 GB at 4-bit
- stablelm-base-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- stablelm-tuned-alpha-7bStability AI7.0B≈4.5 GB at 4-bit
- vicuna-7b-v1.5Large Model Systems Organization7.0B≈4.5 GB at 4-bit
- xgen-7b-8k-baseSalesforce7.0B≈4.5 GB at 4-bit
- granite-4.0-h-tinyIBM Granite6.9B≈4.5 GB at 4-bit
- NuminaMath-7B-TIRProject-Numina6.9B≈4.5 GB at 4-bit
- OLMo-7BAi26.9B≈4.5 GB at 4-bit
- meditron-7bEPFL LLM Team6.7B≈4.4 GB at 4-bit
- Llama-2-7b-chat-hfMeta Llama6.7B≈4.4 GB at 4-bit
- Llama-2-7b-hfMeta Llama6.7B≈4.4 GB at 4-bit
- llama-7bhuggyllama6.7B≈4.4 GB at 4-bit
- llama2_7b_chat_uncensoredgeorgesung6.7B≈4.4 GB at 4-bit
- LlamaGuard-7bMeta Llama6.7B≈4.4 GB at 4-bit1 also selling it hosted
- YuE-s1-7B-anneal-en-cotMultimodal Art Projection6.2B≈4.0 GB at 4-bit
- Yi-6B01-ai6.1B≈3.9 GB at 4-bit1 also selling it hosted
- Yi-6B-200K01-ai6.0B≈3.9 GB at 4-bit
- gpt-j-6bEleutherAI6.0B≈3.9 GB at 4-bit
- pygmalion-6bPygmalion6.0B≈3.9 GB at 4-bit
- DeciLM-6bDeci AI5.7B≈3.7 GB at 4-bit
- Mage-VLMicrosoft4.7B≈3.1 GB at 4-bit
- Qwen-Drive-1.0-4BAlibaba4.5B≈3.0 GB at 4-bit
- Llama-3.1-Minitron-4B-Width-BaseNVIDIA4.5B≈2.9 GB at 4-bit
- medgemma-1.5-4b-itGoogle4.3B≈2.8 GB at 4-bit
- medgemma-4b-itGoogle4.3B≈2.8 GB at 4-bit
- NeoHorse-1-4BTokenRhythm4.2B≈2.7 GB at 4-bit
- Nanbeige4.2-3BNanbeige LLM Lab4.2B≈2.7 GB at 4-bit
- Phi-3-vision-128k-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Phi-3.5-vision-instructMicrosoft4.1B≈2.7 GB at 4-bit1 also selling it hosted
- Spark-X2.5-4BSparkLLM4.1B≈2.7 GB at 4-bit
- MiniCPM-V-4OpenBMB4.1B≈2.6 GB at 4-bit
- AgentCPM-ExploreOpenBMB4.0B≈2.6 GB at 4-bit
- Jan-nano-128kMenlo Research4.0B≈2.6 GB at 4-bit
- Jan-v1-4BJan4.0B≈2.6 GB at 4-bit
- Mellum-4b-baseJetBrains4.0B≈2.6 GB at 4-bit
- GELab-Zero-4B-previewstepfun-ai4.0B≈2.6 GB at 4-bit
- Gemma-3-Gaia-PT-BR-4b-itCEIA-UFG4.0B≈2.6 GB at 4-bit
- LocoTrainer-4BLocoreMind4.0B≈2.6 GB at 4-bit
- MiniCPM3-4BOpenBMB4.0B≈2.6 GB at 4-bit
- Nemotron-Mini-4B-InstructNVIDIA4.0B≈2.6 GB at 4-bit
- Qwen3-4b-Z-Image-Engineer-V4BennyDaBall4.0B≈2.6 GB at 4-bit
- Qwen3-VL-4B-InstructAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- R-4BYannQi4.0B≈2.6 GB at 4-bit
- Youtu-VL-4B-InstructTencent Hunyuan4.0B≈2.6 GB at 4-bit
- gemma-3-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- medgemma-4b-ptGoogle4.0B≈2.6 GB at 4-bit
- t5gemma-2-4b-4bGoogle4.0B≈2.6 GB at 4-bit
- Nanbeige4.1-3BNanbeige LLM Lab3.9B≈2.6 GB at 4-bit
- Phi-4-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Phi-3-mini-128k-instructMicrosoft3.8B≈2.5 GB at 4-bit2 also selling it hosted
- Phi-3-mini-4k-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Phi-3.5-mini-instructMicrosoft3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Hy-Embodied-0.5Tencent Hunyuan3.8B≈2.5 GB at 4-bit
- blip2-opt-2.7bSalesforce3.7B≈2.4 GB at 4-bit
- HyperCLOVAX-SEED-Vision-Instruct-3BHyperCLOVA X3.7B≈2.4 GB at 4-bit
- deepseek-vl2-tinydeepseek-ai3.4B≈2.2 GB at 4-bit
- nllb-200-3.3BAI at Meta3.3B≈2.1 GB at 4-bit
- Llama 3.2 3B InstructMeta3.2B≈2.1 GB at 4-bit14 also selling it hosted
- llama-3.2-Korean-Bllossom-3BBllossom3.2B≈2.1 GB at 4-bit
- imp-v1-3bMILVLG3.2B≈2.1 GB at 4-bit
- LFM2.5-VL-3BLiquid AI3.1B≈2.0 GB at 4-bit
- Qwen2.5-3BAlibaba3.1B≈2.0 GB at 4-bit
- Qwen2.5-3B-InstructAlibaba3.1B≈2.0 GB at 4-bit
- VibeThinker-3BWeiboAI3.1B≈2.0 GB at 4-bit
- SmolLM3-3BHugging Face Smol Models Research3.1B≈2.0 GB at 4-bit
- OpenELM-3B-InstructApple3.0B≈2.0 GB at 4-bit
- SmolLM3-3B-BaseHuggingFaceTB3.0B≈2.0 GB at 4-bit
- bitnet_b1_58-3B1bitLLM3.0B≈2.0 GB at 4-bit
- open_llama_3bOpenLM Research3.0B≈2.0 GB at 4-bit
- open_llama_3b_v2OpenLM Research3.0B≈2.0 GB at 4-bit
- orca_mini_3bpankajmathur3.0B≈2.0 GB at 4-bit
- paligemma2-3b-pt-224Google3.0B≈2.0 GB at 4-bit
- proxy-lite-3bconvergence-ai3.0B≈2.0 GB at 4-bit
- stablelm-3b-4e1tStability AI3.0B≈2.0 GB at 4-bit
- paligemma-3b-pt-224Google2.9B≈1.9 GB at 4-bit
- stablelm-zephyr-3bStability AI2.8B≈1.8 GB at 4-bit
- dolphin-2_6-phi-2Dolphin2.8B≈1.8 GB at 4-bit
- phi-2Microsoft2.8B≈1.8 GB at 4-bit
- Solidity-LLMChainGPT2.8B≈1.8 GB at 4-bit
- gpt-neo-2.7BEleutherAI2.7B≈1.8 GB at 4-bit
- gemma-2-2bGoogle2.6B≈1.7 GB at 4-bit
- gemma-2-2b-itGoogle2.6B≈1.7 GB at 4-bit
- gemma-2-2b-jpn-itGoogle2.6B≈1.7 GB at 4-bit
- LFM2-2.6B-TranscriptLiquid AI2.6B≈1.7 GB at 4-bit