Pass IndexThe State of AISign in

Search models that run on 24 GB

7 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.

Models, tools and agents that go and look something up — the open web, a set of documents, a grounded answer with citations. Pricing is usually per call or per result rather than per token, so the cost follows how often you ask. What differs is freshness, how much of each page you get back, and whether you are handed sources you can check or a summary you must trust.

Wider

Search modelsModels that run on 24 GB