Search models that run on 128 GB
8 models with published weights that fit in 128 GB — a large Mac or a server card. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 136 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A large Mac or a server card. This is where the biggest openly published models come within reach — and where the distinction between total and active parameters starts to decide everything, because a mixture of experts computes with a fraction of itself and must still be held whole.
Models, tools and agents that go and look something up — the open web, a set of documents, a grounded answer with citations. Pricing is usually per call or per result rather than per token, so the cost follows how often you ask. What differs is freshness, how much of each page you get back, and whether you are handed sources you can check or a summary you must trust.
- ERNIE-4.5-VL-28B-A3B-ThinkingBaidu28.0B≈18.2 GB at 4-bit1 also selling it hosted
- Gemma-3-R1984-12BVIDraft12.2B≈7.9 GB at 4-bit
- Molmo2-8BAllen Institute for AI (Ai2)8.7B≈5.6 GB at 4-bit
- Qwen2.5-Coder-7BAlibaba7.0B≈4.5 GB at 4-bit2 also selling it hosted
- Jan-nanoMenlo Research4.0B≈2.6 GB at 4-bit
- LocateAnything-3BNVIDIA3.8B≈2.5 GB at 4-bit
- Isaac 0.2 2B PreviewPerceptron2.0B≈1.3 GB at 4-bit1 also selling it hosted
- Isaac 0.2 1BPerceptron1.0B≈0.7 GB at 4-bit1 also selling it hosted