Embeddings models that run on 16 GB
59 models with published weights that fit in 16 GB — the common laptop. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 16 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
The commonest laptop configuration, and the point where a genuinely useful model runs locally: fourteen billion parameters at four-bit fits with room to spare for the context. Expect a capable general assistant that will not match the frontier, and remember that anything else the machine is doing competes for the same memory.
Models that turn text into a vector so it can be searched by meaning rather than by words. They are cheap — usually cents per million tokens, and often priced for input only, since nothing comes back but numbers. Two things decide the choice: the dimension of the vector, which sets what your database will cost to hold, and whether the model was trained for your language. Changing model later means re-embedding everything you have.
- F2LLM-v2-14Bcodefuse-ai14.0B≈9.1 GB at 4-bit
- KaLM-Embedding-Gemma3-12B-2511Tencent Hunyuan12.0B≈7.8 GB at 4-bit
- bge-multilingual-gemma2BAAI9.2B≈6.0 GB at 4-bit
- F2LLM-v2-8Bcodefuse-ai8.0B≈5.2 GB at 4-bit
- Octen-Embedding-8BOcten8.0B≈5.2 GB at 4-bit
- llama-embed-nemotron-8bNVIDIA8.0B≈5.2 GB at 4-bit
- NV-Embed-v2NVIDIA7.9B≈5.1 GB at 4-bit
- GritLM-7BGritLM7.0B≈4.5 GB at 4-bit
- e5-mistral-7b-instructintfloat7.0B≈4.5 GB at 4-bit
- gte-Qwen2-7B-instructAlibaba-NLP7.0B≈4.5 GB at 4-bit
- speed-embedding-7b-instructHaon-Chen7.0B≈4.5 GB at 4-bit
- dinov3-vit7b16-pretrain-lvd1689mfacebook6.7B≈4.4 GB at 4-bit
- F2LLM-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- F2LLM-v2-4Bcodefuse-ai4.0B≈2.6 GB at 4-bit
- BidirLM-Omni-2.5B-EmbeddingBidirLM2.5B≈1.6 GB at 4-bit
- gte-Qwen2-1.5B-instructAlibaba-NLP1.8B≈1.2 GB at 4-bit
- BidirLM-1.7B-EmbeddingBidirLM1.7B≈1.1 GB at 4-bit
- F2LLM-1.7Bcodefuse-ai1.7B≈1.1 GB at 4-bit
- F2LLM-v2-1.7Bcodefuse-ai1.7B≈1.1 GB at 4-bit
- stella_en_1.5B_v5NovaSearch1.5B≈1.0 GB at 4-bit
- BidirLM-1B-EmbeddingBidirLM1.0B≈0.7 GB at 4-bit
- Nemotron-3-Embed-1B-BF16NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- Nemotron-3-Embed-1B-NVFP4NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- llama-nemotron-embed-vl-1b-v2NVIDIA1.0B≈0.7 GB at 4-bit1 also selling it hosted
- F2LLM-0.6Bcodefuse-ai0.6B≈0.4 GB at 4-bit
- F2LLM-v2-0.6Bcodefuse-ai0.6B≈0.4 GB at 4-bit
- harrier-oss-v1-0.6bMicrosoft0.6B≈0.4 GB at 4-bit
- w2v-bert-2.0facebook0.6B≈0.4 GB at 4-bit
- jina-embeddings-v3Jina AI0.6B≈0.4 GB at 4-bit
- Multilingual-e5-large-instructintfloat0.6B≈0.4 GB at 4-bit4 also selling it hosted
- nomic-embed-text-v2-moeNomic AI0.5B≈0.3 GB at 4-bit
- LaBSEsentence-transformers0.5B≈0.3 GB at 4-bit
- stella_en_400M_v5NovaSearch0.4B≈0.3 GB at 4-bit
- gte-large-en-v1.5Alibaba-NLP0.4B≈0.3 GB at 4-bit
- bge-large-enBAAI0.3B≈0.2 GB at 4-bit1 also selling it hosted
- UAE-Large-V1WhereIsAI0.3B≈0.2 GB at 4-bit
- mxbai-embed-large-v1Mixedbread0.3B≈0.2 GB at 4-bit
- bge-large-zhBAAI0.3B≈0.2 GB at 4-bit
- text2vec-large-chineseGanymedeNil0.3B≈0.2 GB at 4-bit
- gte-multilingual-baseAlibaba-NLP0.3B≈0.2 GB at 4-bit
- dinov3-vitl16-pretrain-lvd1689mAI at Meta0.3B≈0.2 GB at 4-bit
- multilingual-e5-baseintfloat0.3B≈0.2 GB at 4-bit
- paraphrase-multilingual-mpnet-base-v2sentence-transformers0.3B≈0.2 GB at 4-bit
- NeoBERTChandar Research Lab0.2B≈0.2 GB at 4-bit
- jina-clip-v1Jina AI0.2B≈0.1 GB at 4-bit
- jina-embeddings-v2-base-zhJina AI0.2B≈0.1 GB at 4-bit
- gte-modernbert-baseAlibaba-NLP0.1B≈0.1 GB at 4-bit
- modernbert-embed-baseNomic AI0.1B≈0.1 GB at 4-bit
- bart-basefacebook0.1B≈0.1 GB at 4-bit
- jina-embeddings-v2-base-enJina AI0.1B≈0.1 GB at 4-bit
- nomic-embed-text-v1Nomic AI0.1B≈0.1 GB at 4-bit
- distiluse-base-multilingual-cased-v2sentence-transformers0.1B≈0.1 GB at 4-bit
- paraphrase-multilingual-MiniLM-L12-v2sentence-transformers0.1B≈0.1 GB at 4-bit
- pubmedbert-base-embeddingsNeuML0.1B≈0.1 GB at 4-bit
- nomic-embed-vision-v1.5Nomic AI0.1B≈0.1 GB at 4-bit
- dinov2-basefacebook0.1B≈0.1 GB at 4-bit
- vit-base-patch16-224-in21kGoogle0.1B≈0.1 GB at 4-bit
- gte-smallthenlper0.0B≈0.0 GB at 4-bit
- dinov3-vits16-pretrain-lvd1689mfacebook0.0B≈0.0 GB at 4-bit