Pass IndexThe State of AISign in

Reranking search results models that run on 36 GB

10 models with published weights that fit in 36 GB — a MacBook Pro with 36. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 37 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

An Apple configuration, and a comfortable one: it holds what 32 GB holds without the machine feeling tight, which in practice means a longer context or a browser you do not have to close first.

Models that take a handful of results a search already found and put them in the right order. They are the cheap second stage of retrieval: an embedding search casts a wide net fast, a reranker reads the candidates properly. Because they only ever see a few documents, they cost little per query, and they usually improve a weak search more than a better embedding model would.

Wider

Reranking search results modelsModels that run on 36 GB