Crawling the web models that run on 96 GB
2 models with published weights that fit in 96 GB — a MacBook Pro with 96. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 102 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
An Apple configuration for people who intend to run models rather than occasionally try one. Comfortably holds the seventy-billion class with a very long context, or a mixture-of-experts model whose active parameters are few but whose weights are all resident.
Tools that fetch pages and hand them back as something a model can read: rendered HTML, markdown, a screenshot, a whole site map. Priced per page or per session. The differences that matter are whether JavaScript is executed, what happens at a login or a bot check, and how gracefully the thing fails on a site that does not want to be read.
- xlm-roberta-largeFacebook AI community0.6B≈0.4 GB at 4-bit
- xlm-roberta-baseFacebook AI community0.3B≈0.2 GB at 4-bit