Safety classification models that run on 96 GB
4 models with published weights that fit in 96 GB — a MacBook Pro with 96. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 102 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
An Apple configuration for people who intend to run models rather than occasionally try one. Comfortably holds the seventy-billion class with a very long context, or a mixture-of-experts model whose active parameters are few but whose weights are all resident.
Models that classify content rather than produce it: what is unsafe, what breaks a policy, what should not be shown. They sit in front of or behind a bigger model, are small and cheap by design, and are the part of a system nobody notices until it is wrong. The thing to check is what the model was trained to catch, because policies differ and a guard tuned for one product's rules will be strict in the wrong places for yours.
- Llama Guard 4 12BMeta12.0B≈7.8 GB at 4-bit6 also selling it hosted
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Llama-Guard-3-8BMeta8.0B≈5.2 GB at 4-bit4 also selling it hosted
- shieldgemma-2-4b-itGoogle4.0B≈2.6 GB at 4-bit