Safety classification models that run on 128 GB
5 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.
Models that classify content rather than produce it: what is unsafe, what breaks a policy, what should not be shown. They sit in front of or behind a bigger model, are small and cheap by design, and are the part of a system nobody notices until it is wrong. The thing to check is what the model was trained to catch, because policies differ and a guard tuned for one product's rules will be strict in the wrong places for yours.
- GPT OSS Safeguard 120BOpenAI120B≈78.0 GB at 4-bit3 also selling it hosted
- Llama Guard 4 12BMeta12.0B≈7.8 GB at 4-bit6 also selling it hosted
- Llama Guard 3 11B VisionMeta11.0B≈7.2 GB at 4-bit2 also selling it hosted
- Llama-Guard-3-8BMeta8.0B≈5.2 GB at 4-bit4 also selling it hosted
- shieldgemma-2-4b-itGoogle4.0B≈2.6 GB at 4-bit