Translation models that run on 24 GB
16 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
Models and services aimed at moving text between languages. A general chat model will translate too, and often well; a dedicated one earns its place with glossaries, formality control, document formats it keeps intact, and rates that assume volume. Check which direction was measured — most quality figures are for translating into English, and the other direction is harder.
- Voxtral Small 24B 2507Mistral AI24.3B≈15.8 GB at 4-bit4 also selling it hosted
- translategemma-12b-itGoogle13.2B≈8.6 GB at 4-bit
- granite-speech-3.3-8bIBM Granite8.0B≈5.2 GB at 4-bit
- Hy-MT2-7BTencent Hunyuan7.0B≈4.5 GB at 4-bit1 also selling it hosted
- Seed-X-PPO-7BByteDance Seed7.0B≈4.5 GB at 4-bit
- translategemma-4b-itGoogle5.0B≈3.2 GB at 4-bit
- Voxtral-Mini-3B-2507Mistral AI3.0B≈2.0 GB at 4-bit2 also selling it hosted
- madlad400-3b-mtGoogle2.9B≈1.9 GB at 4-bit
- seamless-m4t-v2-largeAI at Meta2.3B≈1.5 GB at 4-bit
- Hy-MT2-1.8BTencent Hunyuan1.8B≈1.2 GB at 4-bit1 also selling it hosted
- Whisper Large v3OpenAI1.5B≈1.0 GB at 4-bit2 also selling it hosted
- whisper-largeOpenAI1.5B≈1.0 GB at 4-bit
- whisper-large-v2OpenAI1.5B≈1.0 GB at 4-bit
- canary-1b-v2NVIDIA1.0B≈0.6 GB at 4-bit
- whisper-smallOpenAI0.2B≈0.2 GB at 4-bit
- whisper-tinyOpenAI0.0B≈0.0 GB at 4-bit