Reasoning models that run on 256 GB
86 models with published weights that fit in 256 GB — a Mac Studio with 256. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 274 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A Mac Studio at the top of its configuration, or a serious machine. Almost everything openly published fits, including the large mixtures of experts. At this point the constraint is no longer whether the model loads but whether it generates fast enough to be worth waiting for.
Models that plan before they answer, spending extra tokens on working the problem through. They win on mathematics, hard code and anything with several steps, and they are measured on boards like GPQA, AIME and ARC-AGI. The catch is what the thinking costs: reasoning is billed as output tokens, so the same model can cost several times more per answer at a high effort than a low one, and for a question that needed one turn you have paid for a monologue.
- Inkling SmallThinking Machines Lab266B≈172.9 GB at 4-bit8 also selling it hosted
- Qwen3 VL 235B A22B ThinkingAlibaba236B≈153.2 GB at 4-bit9 also selling it hosted
- Qwen3 235B A22BAlibaba235B≈152.8 GB at 4-bit10 also selling it hosted
- Qwen3 235B A22B Thinking 2507Alibaba235B≈152.8 GB at 4-bit12 also selling it hosted
- Qwen3-235B-A22B-Instruct-2507Alibaba235B≈152.8 GB at 4-bit13 also selling it hosted
- MiniMax M2.5MiniMax229B≈148.7 GB at 4-bit19 also selling it hosted
- MiniMax M2.1MiniMax229B≈148.6 GB at 4-bit11 also selling it hosted
- Step 3.7 FlashStepFun201B≈130.9 GB at 4-bit9 also selling it hosted
- Step 3.5 FlashStepFun199B≈129.6 GB at 4-bit8 also selling it hosted
- WizardLM-2 8x22BMicrosoft141B≈91.4 GB at 4-bit6 also selling it hosted
- Qwen3.5-122B-A10BAlibaba125B≈81.3 GB at 4-bit8 also selling it hosted
- Nemotron 3 SuperNVIDIA124B≈80.3 GB at 4-bit7 also selling it hosted
- GPT OSS 120BOpenAI117B≈75.9 GB at 4-bit29 also selling it hosted
- GLM-4.5-AirZ.ai110B≈71.8 GB at 4-bit12 also selling it hosted
- Llama-3.2-90B-Vision-Instruct90.0B≈58.5 GB at 4-bit7 also selling it hosted
- Qwen3 Next 80B A3B ThinkingAlibaba81.3B≈52.9 GB at 4-bit9 also selling it hosted
- Qwen2.5 72B InstructAlibaba72.7B≈47.3 GB at 4-bit12 also selling it hosted
- Qwen1.5-72B-ChatAlibaba72.3B≈47.0 GB at 4-bit1 also selling it hosted
- Qwen2-72B-InstructAlibaba72.0B≈46.8 GB at 4-bit3 also selling it hosted
- Llama 3.3 70B InstructMeta70.6B≈45.9 GB at 4-bit29 also selling it hosted
- R1 Distill Llama 70BDeepSeek70.6B≈45.9 GB at 4-bit12 also selling it hosted
- Llama-2-70b-chat-hfMeta70.0B≈45.5 GB at 4-bit1 also selling it hosted
- Meta-Llama-3-70B-InstructMeta70.0B≈45.5 GB at 4-bit6 also selling it hosted
- Meta-Llama-3.1-70B-InstructMeta70.0B≈45.5 GB at 4-bit7 also selling it hosted
- deepseek-llm-67b-chatdeepseek-ai67.0B≈43.6 GB at 4-bit
- Qwen3.5-40B-Claude-4.6-Opus-Deckard-Heretic-Uncensored-ThinkingDavidAU39.5B≈25.7 GB at 4-bit
- Seed-OSS-36B-InstructByteDance Seed36.2B≈23.5 GB at 4-bit1 also selling it hosted
- Qwen3.5-35B-A3BAlibaba36.0B≈23.4 GB at 4-bit9 also selling it hosted
- Qwen3.6 35B A3BAlibaba36.0B≈23.4 GB at 4-bit10 also selling it hosted
- Yi-34B-Chat01-ai34.4B≈22.4 GB at 4-bit1 also selling it hosted
- Yi-1.5-34B-Chat01-ai34.0B≈22.1 GB at 4-bit
- Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16NVIDIA33.0B≈21.5 GB at 4-bit
- Qwen3 32BAlibaba32.8B≈21.3 GB at 4-bit17 also selling it hosted
- AM-Thinking-v1am team32.8B≈21.3 GB at 4-bit
- DeepSeek R1 Distill QWEN 32BDeepSeek32.0B≈20.8 GB at 4-bit6 also selling it hosted
- OpenReasoning-Nemotron-32BNVIDIA32.0B≈20.8 GB at 4-bit1 also selling it hosted
- QwQ-32BAlibaba32.0B≈20.8 GB at 4-bit7 also selling it hosted
- Gemma 4 31BGoogle31.3B≈20.3 GB at 4-bit15 also selling it hosted
- Qwen3 VL 30B A3B ThinkingAlibaba31.1B≈20.2 GB at 4-bit7 also selling it hosted
- Qwen3 30B A3BAlibaba30.5B≈19.8 GB at 4-bit11 also selling it hosted
- Qwen3 30B A3B Instruct 2507Alibaba30.5B≈19.8 GB at 4-bit8 also selling it hosted
- Qwen3 30B A3B Thinking 2507Alibaba30.5B≈19.8 GB at 4-bit4 also selling it hosted
- Qwen3 Omni 30B A3B ThinkingAlibaba30.0B≈19.5 GB at 4-bit2 also selling it hosted
- ERNIE-4.5-VL-28B-A3B-ThinkingBaidu28.0B≈18.2 GB at 4-bit1 also selling it hosted
- Qwen3.5-27B-Claude-4.6-Opus-Reasoning-DistilledJackrong27.8B≈18.1 GB at 4-bit
- Qwen3.6 27BAlibaba27.8B≈18.1 GB at 4-bit12 also selling it hosted
- Gemma 3 27BGoogle27.4B≈17.8 GB at 4-bit11 also selling it hosted
- Gemma 2 27BGoogle27.2B≈17.7 GB at 4-bit4 also selling it hosted
- thinkingcap-qwen3.6-27bsference27.0B≈17.6 GB at 4-bit1 also selling it hosted
- ERNIE-4.5-21B-A3B-ThinkingBaidu21.0B≈13.7 GB at 4-bit2 also selling it hosted
- GPT OSS 20BOpenAI20.9B≈13.6 GB at 4-bit20 also selling it hosted
- Kimi-VL-A3B-ThinkingMoonshot AI16.4B≈10.7 GB at 4-bit
- Kimi-VL-A3B-Thinking-2506Moonshot AI16.4B≈10.7 GB at 4-bit
- Phi-4-reasoning-vision-15BMicrosoft15.0B≈9.8 GB at 4-bit
- Qwen3 14BAlibaba14.8B≈9.6 GB at 4-bit8 also selling it hosted
- Phi-4Microsoft14.7B≈9.5 GB at 4-bit6 also selling it hosted
- DeepSeek R1 Distill QWEN 14BDeepSeek14.0B≈9.1 GB at 4-bit5 also selling it hosted
- Phi-3-medium-128k-instructMicrosoft14.0B≈9.1 GB at 4-bit1 also selling it hosted
- Gemma 3 12BGoogle12.2B≈7.9 GB at 4-bit9 also selling it hosted
- Mellum2-12B-A2.5B-ThinkingJetBrains12.1B≈7.9 GB at 4-bit
- Qwen3.5 9BAlibaba9.7B≈6.3 GB at 4-bit8 also selling it hosted
- Qwen3.5-9B-Claude-4.6-HighIQ-THINKING-HERETIC-UNCENSOREDDavidAU9.4B≈6.1 GB at 4-bit
- gemma-2-9b-itGoogle9.0B≈5.9 GB at 4-bit3 also selling it hosted
- Qwen3 VL 8B ThinkingAlibaba8.8B≈5.7 GB at 4-bit3 also selling it hosted
- NuMarkdown-8B-ThinkingNuMind8.3B≈5.4 GB at 4-bit
- Qwen3 8BAlibaba8.2B≈5.3 GB at 4-bit8 also selling it hosted
- Llama 3.1 8B InstructMeta8.0B≈5.2 GB at 4-bit20 also selling it hosted
- DeepSeek R1 0528 Qwen3 8BDeepSeek8.0B≈5.2 GB at 4-bit1 also selling it hosted
- Seed-Coder-8B-ReasoningByteDance Seed8.0B≈5.2 GB at 4-bit
- Qwen2.5 7B InstructAlibaba7.6B≈5.0 GB at 4-bit8 also selling it hosted
- Mistral-7B-Instruct-v0.3Mistral AI7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Gemma 3 4BGoogle4.3B≈2.8 GB at 4-bit6 also selling it hosted
- DASD-4B-ThinkingAlibaba Cloud Apsara Lab4.0B≈2.6 GB at 4-bit
- Qwen3 4BAlibaba4.0B≈2.6 GB at 4-bit3 also selling it hosted
- Qwen3-4B-Instruct-2507Alibaba4.0B≈2.6 GB at 4-bit2 also selling it hosted
- Qwen3-4B-Thinking-2507Alibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3-VL-4B-ThinkingAlibaba4.0B≈2.6 GB at 4-bit1 also selling it hosted
- Qwen3.5-4BAlibaba4.0B≈2.6 GB at 4-bit2 also selling it hosted
- Granite 4.0 MicroIBM watsonx.ai3.2B≈2.1 GB at 4-bit2 also selling it hosted
- Ouro-2.6B-ThinkingByteDance2.6B≈1.7 GB at 4-bit
- Qwen3-1.7BAlibaba2.0B≈1.3 GB at 4-bit1 also selling it hosted
- DeepSeek-R1-Distill-Qwen-1.5BDeepSeek1.8B≈1.2 GB at 4-bit3 also selling it hosted
- Nemotron-Research-Reasoning-Qwen-1.5BNVIDIA1.8B≈1.2 GB at 4-bit
- Llama 3.2 1B InstructMeta1.2B≈0.8 GB at 4-bit11 also selling it hosted
- LFM2.5-1.2B-ThinkingLiquid AI1.2B≈0.8 GB at 4-bit
- MiniCPM5-1B-Claude-Opus-Fable5-ThinkingGnLOLot1.0B≈0.7 GB at 4-bit