Crawling the web models that run on 24 GB
2 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
Tools that fetch pages and hand them back as something a model can read: rendered HTML, markdown, a screenshot, a whole site map. Priced per page or per session. The differences that matter are whether JavaScript is executed, what happens at a login or a bot check, and how gracefully the thing fails on a site that does not want to be read.
- xlm-roberta-largeFacebook AI community0.6B≈0.4 GB at 4-bit
- xlm-roberta-baseFacebook AI community0.3B≈0.2 GB at 4-bit