Pass IndexThe State of AISign in

Embeddings models that run on 32 GB

60 models with published weights that fit in 32 GB — a well-specified laptop. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 33 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

Enough for a thirty-billion-parameter model at four-bit with a long context, or a smaller one at higher precision if quality matters more than size. A practical ceiling for a laptop that also has to be a laptop.

Models that turn text into a vector so it can be searched by meaning rather than by words. They are cheap — usually cents per million tokens, and often priced for input only, since nothing comes back but numbers. Two things decide the choice: the dimension of the vector, which sets what your database will cost to hold, and whether the model was trained for your language. Changing model later means re-embedding everything you have.

Wider

Embeddings modelsModels that run on 32 GB