Evaluation models that run on 24 GB
14 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
Tools for judging what a model produced: scoring runs, tracing a chain of calls, catching a regression before a user does. Priced per trace or per evaluation. Most use a model as the judge, which is worth knowing, because it means your evaluation has a bill and an opinion of its own.
- DeepSeek-Coder-V2-Lite-InstructDeepSeek15.7B≈10.2 GB at 4-bit1 also selling it hosted
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Qwythos-9B-Claude-Mythos-5-1Mempero-ai9.4B≈6.1 GB at 4-bit
- Nemotron Nano 9B V2NVIDIA9.0B≈5.9 GB at 4-bit5 also selling it hosted
- Aion-RP 1.0 (8B)AionLabs8.0B≈5.2 GB at 4-bit2 also selling it hosted
- DeepSeek R1 0528 Qwen3 8BDeepSeek8.0B≈5.2 GB at 4-bit1 also selling it hosted
- DeepSeek-R1-Distill-Llama-8BDeepSeek8.0B≈5.2 GB at 4-bit3 also selling it hosted
- Qwen2.5-Omni-3BAlibaba5.5B≈3.6 GB at 4-bit
- fable-tracesAliesTaha4.0B≈2.6 GB at 4-bit
- LocoOperator-4BLocoreMind4.0B≈2.6 GB at 4-bit
- Isaac 0.2 2B PreviewPerceptron2.0B≈1.3 GB at 4-bit1 also selling it hosted
- Isaac 0.2 1BPerceptron1.0B≈0.7 GB at 4-bit1 also selling it hosted
- llama-nemotron-rerank-vl-1b-v2NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- PaddleOCR-VL-0.9Bpaddlepaddle0.9B≈0.6 GB at 4-bit1 also selling it hosted