Safety classification models that run on 36 GB
4 models with published weights that fit in 36 GB — a MacBook Pro with 36. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 37 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
An Apple configuration, and a comfortable one: it holds what 32 GB holds without the machine feeling tight, which in practice means a longer context or a browser you do not have to close first.
Models that classify content rather than produce it: what is unsafe, what breaks a policy, what should not be shown. They sit in front of or behind a bigger model, are small and cheap by design, and are the part of a system nobody notices until it is wrong. The thing to check is what the model was trained to catch, because policies differ and a guard tuned for one product's rules will be strict in the wrong places for yours.
- Llama Guard 4 12BMeta12.0B≈7.8 GB at 4-bit6 also selling it hosted
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Llama-Guard-3-8BMeta8.0B≈5.2 GB at 4-bit4 also selling it hosted
- shieldgemma-2-4b-itGoogle4.0B≈2.6 GB at 4-bit