Safety classification models that run on 24 GB
4 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
Models that classify content rather than produce it: what is unsafe, what breaks a policy, what should not be shown. They sit in front of or behind a bigger model, are small and cheap by design, and are the part of a system nobody notices until it is wrong. The thing to check is what the model was trained to catch, because policies differ and a guard tuned for one product's rules will be strict in the wrong places for yours.
- Llama Guard 4 12BMeta12.0B≈7.8 GB at 4-bit6 also selling it hosted
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Llama-Guard-3-8BMeta8.0B≈5.2 GB at 4-bit4 also selling it hosted
- shieldgemma-2-4b-itGoogle4.0B≈2.6 GB at 4-bit