Pass IndexThe State of AISign in

Embeddings models that run on 64 GB

60 models with published weights that fit in 64 GB — a workstation. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 67 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.

A workstation, or a well-specified Mac. Seventy-billion-parameter models fit at four-bit, which is the class where local output stops being obviously worse than the hosted models people pay for. Loading takes real time and the machine will be warm.

Models that turn text into a vector so it can be searched by meaning rather than by words. They are cheap — usually cents per million tokens, and often priced for input only, since nothing comes back but numbers. Two things decide the choice: the dimension of the vector, which sets what your database will cost to hold, and whether the model was trained for your language. Changing model later means re-embedding everything you have.

Wider

Embeddings modelsModels that run on 64 GB