Embeddings models that run on 256 GB
60 models with published weights that fit in 256 GB — a Mac Studio with 256. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 274 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A Mac Studio at the top of its configuration, or a serious machine. Almost everything openly published fits, including the large mixtures of experts. At this point the constraint is no longer whether the model loads but whether it generates fast enough to be worth waiting for.
Models that turn text into a vector so it can be searched by meaning rather than by words. They are cheap — usually cents per million tokens, and often priced for input only, since nothing comes back but numbers. Two things decide the choice: the dimension of the vector, which sets what your database will cost to hold, and whether the model was trained for your language. Changing model later means re-embedding everything you have.
- harrier-oss-v1-27bMicrosoft27.0B≈17.6 GB at 4-bit
- F2LLM-v2-14Bcodefuse-ai14.0B≈9.1 GB at 4-bit
- KaLM-Embedding-Gemma3-12B-2511Tencent Hunyuan12.0B≈7.8 GB at 4-bit
- bge-multilingual-gemma2BAAI9.2B≈6.0 GB at 4-bit
- F2LLM-v2-8Bcodefuse-ai8.0B≈5.2 GB at 4-bit
- Octen-Embedding-8BOcten8.0B≈5.2 GB at 4-bit
- llama-embed-nemotron-8bNVIDIA8.0B≈5.2 GB at 4-bit
- NV-Embed-v2NVIDIA7.9B≈5.1 GB at 4-bit
- GritLM-7BGritLM7.0B≈4.5 GB at 4-bit
- e5-mistral-7b-instructintfloat7.0B≈4.5 GB at 4-bit
- gte-Qwen2-7B-instructAlibaba-NLP7.0B≈4.5 GB at 4-bit
- speed-embedding-7b-instructHaon-Chen7.0B≈4.5 GB at 4-bit
- dinov3-vit7b16-pretrain-lvd1689mfacebook6.7B≈4.4 GB at 4-bit
- F2LLM-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- F2LLM-v2-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- BidirLM-Omni-2.5B-EmbeddingBidirLM2.5B≈1.6 GB at 4-bit
- gte-Qwen2-1.5B-instructAlibaba-NLP1.8B≈1.2 GB at 4-bit
- BidirLM-1.7B-EmbeddingBidirLM1.7B≈1.1 GB at 4-bit
- F2LLM-1.7Bcodefuse-ai1.7B≈1.1 GB at 4-bit
- F2LLM-v2-1.7Bcodefuse-ai1.7B≈1.1 GB at 4-bit
- stella_en_1.5B_v5NovaSearch1.5B≈1.0 GB at 4-bit
- BidirLM-1B-EmbeddingBidirLM1.0B≈0.7 GB at 4-bit
- Nemotron-3-Embed-1B-BF16NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- Nemotron-3-Embed-1B-NVFP4NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- llama-nemotron-embed-vl-1b-v2NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- F2LLM-0.6Bcodefuse-ai0.6B≈0.4 GB at 4-bit
- F2LLM-v2-0.6Bcodefuse-ai0.6B≈0.4 GB at 4-bit
- harrier-oss-v1-0.6bMicrosoft0.6B≈0.4 GB at 4-bit
- w2v-bert-2.0facebook0.6B≈0.4 GB at 4-bit
- jina-embeddings-v3Jina AI0.6B≈0.4 GB at 4-bit
- Multilingual-e5-large-instructintfloat0.6B≈0.4 GB at 4-bit4 also selling it hosted
- nomic-embed-text-v2-moeNomic AI0.5B≈0.3 GB at 4-bit
- LaBSEsentence-transformers0.5B≈0.3 GB at 4-bit
- stella_en_400M_v5NovaSearch0.4B≈0.3 GB at 4-bit
- gte-large-en-v1.5Alibaba-NLP0.4B≈0.3 GB at 4-bit
- bge-large-enBAAI0.3B≈0.2 GB at 4-bit1 also selling it hosted
- UAE-Large-V1WhereIsAI0.3B≈0.2 GB at 4-bit
- mxbai-embed-large-v1Mixedbread0.3B≈0.2 GB at 4-bit
- bge-large-zhBAAI0.3B≈0.2 GB at 4-bit
- text2vec-large-chineseGanymedeNil0.3B≈0.2 GB at 4-bit
- gte-multilingual-baseAlibaba-NLP0.3B≈0.2 GB at 4-bit
- dinov3-vitl16-pretrain-lvd1689mAI at Meta0.3B≈0.2 GB at 4-bit
- multilingual-e5-baseintfloat0.3B≈0.2 GB at 4-bit
- paraphrase-multilingual-mpnet-base-v2sentence-transformers0.3B≈0.2 GB at 4-bit
- NeoBERTChandar Research Lab0.2B≈0.2 GB at 4-bit
- jina-clip-v1Jina AI0.2B≈0.1 GB at 4-bit
- jina-embeddings-v2-base-zhJina AI0.2B≈0.1 GB at 4-bit
- gte-modernbert-baseAlibaba-NLP0.1B≈0.1 GB at 4-bit
- modernbert-embed-baseNomic AI0.1B≈0.1 GB at 4-bit
- bart-basefacebook0.1B≈0.1 GB at 4-bit
- jina-embeddings-v2-base-enJina AI0.1B≈0.1 GB at 4-bit
- nomic-embed-text-v1Nomic AI0.1B≈0.1 GB at 4-bit
- distiluse-base-multilingual-cased-v2sentence-transformers0.1B≈0.1 GB at 4-bit
- paraphrase-multilingual-MiniLM-L12-v2sentence-transformers0.1B≈0.1 GB at 4-bit
- pubmedbert-base-embeddingsNeuML0.1B≈0.1 GB at 4-bit
- nomic-embed-vision-v1.5Nomic AI0.1B≈0.1 GB at 4-bit
- dinov2-basefacebook0.1B≈0.1 GB at 4-bit
- vit-base-patch16-224-in21kGoogle0.1B≈0.1 GB at 4-bit
- gte-smallthenlper0.0B≈0.0 GB at 4-bit
- dinov3-vits16-pretrain-lvd1689mfacebook0.0B≈0.0 GB at 4-bit