Translation models that run on 8 GB
13 models with published weights that fit in 8 GB — a phone, a base iPad, an Air. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 7 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
The memory of a phone, a base iPad or an entry-level laptop. What fits is small: models of a few billion parameters, quick and cheap to run, good at summarising, classifying and simple extraction, and out of their depth on long reasoning. This is also where on-device makes the most sense, because the alternative is a network round trip for something that takes a moment.
Models and services aimed at moving text between languages. A general chat model will translate too, and often well; a dedicated one earns its place with glossaries, formality control, document formats it keeps intact, and rates that assume volume. Check which direction was measured — most quality figures are for translating into English, and the other direction is harder.
- Hy-MT2-7BTencent Hunyuan7.0B≈4.5 GB at 4-bit1 also selling it hosted
- Seed-X-PPO-7BByteDance Seed7.0B≈4.5 GB at 4-bit
- translategemma-4b-itGoogle5.0B≈3.2 GB at 4-bit
- Voxtral-Mini-3B-2507Mistral AI3.0B≈2.0 GB at 4-bit2 also selling it hosted
- madlad400-3b-mtGoogle2.9B≈1.9 GB at 4-bit
- seamless-m4t-v2-largeAI at Meta2.3B≈1.5 GB at 4-bit
- Hy-MT2-1.8BTencent Hunyuan1.8B≈1.2 GB at 4-bit1 also selling it hosted
- Whisper Large v3OpenAI1.5B≈1.0 GB at 4-bit2 also selling it hosted
- whisper-largeOpenAI1.5B≈1.0 GB at 4-bit
- whisper-large-v2OpenAI1.5B≈1.0 GB at 4-bit
- canary-1b-v2NVIDIA1.0B≈0.6 GB at 4-bit
- whisper-smallOpenAI0.2B≈0.2 GB at 4-bit
- whisper-tinyOpenAI0.0B≈0.0 GB at 4-bit