Speech models that run on 24 GB
51 models with published weights that fit in 24 GB — a 24 GB card, or a Mac with 24. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 24 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A 24 GB graphics card or a Mac configured with 24. This is the first size where the well-regarded mid-weight models — the twenty-something billion parameter class — run comfortably, and where a local model starts to be a real alternative to an API for daily work rather than a demonstration.
Models that read text aloud. They are usually priced per character, which makes the arithmetic simple: a thousand characters is roughly a paragraph. The real choices are the voice itself, whether you may clone one, how much control you have over emotion and pacing, and latency — a model that sounds wonderful in a rendered file may be too slow to hold a conversation.
- Kimi-Audio-7B-InstructMoonshot AI9.8B≈6.3 GB at 4-bit
- VibeVoice-7BVibeVoice Community (Unofficial)9.3B≈6.1 GB at 4-bit
- VibeVoice-Largeaoi-ot9.3B≈6.1 GB at 4-bit
- kugelaudio-0-openKugelaudio9.3B≈6.1 GB at 4-bit
- MOSS-TTSOpenMOSS-Team8.5B≈5.5 GB at 4-bit
- MOSS-TTS-v1.5OpenMOSS-Team8.5B≈5.5 GB at 4-bit
- personaplex-7b-v1NVIDIA8.4B≈5.4 GB at 4-bit
- MisoTTSMiso Labs8.2B≈5.3 GB at 4-bit
- higgs-tts-2-3b-baseBoson AI5.8B≈3.8 GB at 4-bit
- acestep-v15-xl-turboACE-Step5.0B≈3.2 GB at 4-bit
- higgs-tts-3-4bBoson AI4.7B≈3.0 GB at 4-bit
- s2-proFish Audio4.6B≈3.0 GB at 4-bit1 also selling it hosted
- Llasa-3BHKUST Audio4.0B≈2.6 GB at 4-bit
- Voxtral-4B-TTS-2603Mistral AI4.0B≈2.6 GB at 4-bit
- HeartMuLa-oss-3BHeartMuLa3.9B≈2.6 GB at 4-bit
- Orpheus 3BCanopy Labs3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Step-Audio-TTS-3Bstepfun-ai3.5B≈2.3 GB at 4-bit
- Breeze-TTS-2BreezeBlue3.5B≈2.3 GB at 4-bit
- fish-agent-v0.1-3bFish Audio3.0B≈2.0 GB at 4-bit
- orpheus-3b-0.1-pretrainedCanopy Labs3.0B≈2.0 GB at 4-bit
- tada-3b-mlHume AI3.0B≈2.0 GB at 4-bit
- VibeVoice-1.5BMicrosoft2.7B≈1.8 GB at 4-bit
- VoxCPM2OpenBMB2.3B≈1.5 GB at 4-bit
- tada-1bHume AI2.2B≈1.4 GB at 4-bit
- SoulX-Podcast-1.7BSoul-AILab2.1B≈1.3 GB at 4-bit
- Dia2-2BNari Labs1.9B≈1.2 GB at 4-bit
- Qwen3-TTS-12Hz-1.7B-CustomVoiceAlibaba1.9B≈1.2 GB at 4-bit
- Qwen3-TTS-12Hz-1.7B-VoiceDesignAlibaba1.9B≈1.2 GB at 4-bit
- Dia-1.6BNari Labs1.6B≈1.0 GB at 4-bit
- tts-1.6b-en_frKyutai1.6B≈1.0 GB at 4-bit
- LFM2-Audio-1.5BLiquid AI1.5B≈1.0 GB at 4-bit
- LFM2.5-Audio-1.5BLiquid AI1.5B≈1.0 GB at 4-bit
- Llama-OuteTTS-1.0-1BOuteAI1.2B≈0.8 GB at 4-bit
- stable-audio-open-1.0Stability AI1.2B≈0.8 GB at 4-bit
- VibeVoice-Realtime-0.5BMicrosoft1.0B≈0.7 GB at 4-bit
- csm-1bsesame1.0B≈0.7 GB at 4-bit1 also selling it hosted
- metavoice-1B-v0.1MetaVoice1.0B≈0.7 GB at 4-bit
- indic-parler-ttsAI4Bharat0.9B≈0.6 GB at 4-bit
- Qwen3-TTS-12Hz-0.6B-CustomVoiceAlibaba0.9B≈0.6 GB at 4-bit
- VoxCPM1.5OpenBMB0.8B≈0.5 GB at 4-bit
- neutts-airNeuphonic0.7B≈0.5 GB at 4-bit
- parler_tts_mini_v0.1Parler TTS0.6B≈0.4 GB at 4-bit
- Audio8-TTS-Preview-0.6bAudio80.6B≈0.4 GB at 4-bit1 also selling it hosted
- Qwen3-TTS-12Hz-0.6B-BaseAlibaba0.6B≈0.4 GB at 4-bit
- MiraTTSYatharthS0.5B≈0.3 GB at 4-bit
- Fun-CosyVoice3-0.5B-2512QwenAudio0.5B≈0.3 GB at 4-bit
- VoxCPM-0.5BOpenBMB0.5B≈0.3 GB at 4-bit
- IndicF5AI4Bharat0.4B≈0.2 GB at 4-bit
- VoiceCraftpyp10.3B≈0.2 GB at 4-bit
- Audio8-TTS-Preview-0.1bEdge00.2B≈0.1 GB at 4-bit
- Soprano-1.1-80Mekwek0.1B≈0.1 GB at 4-bit