Safety classification models that run on 64 GB
4 models with published weights that fit in 64 GB — a workstation. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 67 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A workstation, or a well-specified Mac. Seventy-billion-parameter models fit at four-bit, which is the class where local output stops being obviously worse than the hosted models people pay for. Loading takes real time and the machine will be warm.
Models that classify content rather than produce it: what is unsafe, what breaks a policy, what should not be shown. They sit in front of or behind a bigger model, are small and cheap by design, and are the part of a system nobody notices until it is wrong. The thing to check is what the model was trained to catch, because policies differ and a guard tuned for one product's rules will be strict in the wrong places for yours.
- Llama Guard 4 12BMeta12.0B≈7.8 GB at 4-bit6 also selling it hosted
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Llama-Guard-3-8BMeta8.0B≈5.2 GB at 4-bit4 also selling it hosted
- shieldgemma-2-4b-itGoogle4.0B≈2.6 GB at 4-bit