Speech models that run on 256 GB
52 models with published weights that fit in 256 GB — a Mac Studio with 256. Counted at four-bit quantisation, weights only, leaving the machine about a third of its memory and a gigabyte for context: room for roughly 274 billion parameters. A mixture of experts is counted in full, because it is held in full even though only a few experts compute.
A Mac Studio at the top of its configuration, or a serious machine. Almost everything openly published fits, including the large mixtures of experts. At this point the constraint is no longer whether the model loads but whether it generates fast enough to be worth waiting for.
Models that read text aloud. They are usually priced per character, which makes the arithmetic simple: a thousand characters is roughly a paragraph. The real choices are the voice itself, whether you may clone one, how much control you have over emotion and pacing, and latency — a model that sounds wonderful in a rendered file may be too slow to hold a conversation.
- Qwen3 Omni 30B A3B InstructAlibaba30.0B≈19.5 GB at 4-bit2 also selling it hosted
- Kimi-Audio-7B-InstructMoonshot AI9.8B≈6.3 GB at 4-bit
- VibeVoice-7BVibeVoice Community (Unofficial)9.3B≈6.1 GB at 4-bit
- VibeVoice-Largeaoi-ot9.3B≈6.1 GB at 4-bit
- kugelaudio-0-openKugelaudio9.3B≈6.1 GB at 4-bit
- MOSS-TTSOpenMOSS-Team8.5B≈5.5 GB at 4-bit
- MOSS-TTS-v1.5OpenMOSS-Team8.5B≈5.5 GB at 4-bit
- personaplex-7b-v1NVIDIA8.4B≈5.4 GB at 4-bit
- MisoTTSMiso Labs8.2B≈5.3 GB at 4-bit
- higgs-tts-2-3b-baseBoson AI5.8B≈3.8 GB at 4-bit
- acestep-v15-xl-turboACE-Step5.0B≈3.2 GB at 4-bit
- higgs-tts-3-4bBoson AI4.7B≈3.0 GB at 4-bit
- s2-proFish Audio4.6B≈3.0 GB at 4-bit1 also selling it hosted
- Llasa-3BHKUST Audio4.0B≈2.6 GB at 4-bit
- Voxtral-4B-TTS-2603Mistral AI4.0B≈2.6 GB at 4-bit
- HeartMuLa-oss-3BHeartMuLa3.9B≈2.6 GB at 4-bit
- Orpheus 3BCanopy Labs3.8B≈2.5 GB at 4-bit1 also selling it hosted
- Step-Audio-TTS-3Bstepfun-ai3.5B≈2.3 GB at 4-bit
- Breeze-TTS-2BreezeBlue3.5B≈2.3 GB at 4-bit
- fish-agent-v0.1-3bFish Audio3.0B≈2.0 GB at 4-bit
- orpheus-3b-0.1-pretrainedCanopy Labs3.0B≈2.0 GB at 4-bit
- tada-3b-mlHume AI3.0B≈2.0 GB at 4-bit
- VibeVoice-1.5BMicrosoft2.7B≈1.8 GB at 4-bit
- VoxCPM2OpenBMB2.3B≈1.5 GB at 4-bit
- tada-1bHume AI2.2B≈1.4 GB at 4-bit
- SoulX-Podcast-1.7BSoul-AILab2.1B≈1.3 GB at 4-bit
- Dia2-2BNari Labs1.9B≈1.2 GB at 4-bit
- Qwen3-TTS-12Hz-1.7B-CustomVoiceAlibaba1.9B≈1.2 GB at 4-bit
- Qwen3-TTS-12Hz-1.7B-VoiceDesignAlibaba1.9B≈1.2 GB at 4-bit
- Dia-1.6BNari Labs1.6B≈1.0 GB at 4-bit
- tts-1.6b-en_frKyutai1.6B≈1.0 GB at 4-bit
- LFM2-Audio-1.5BLiquid AI1.5B≈1.0 GB at 4-bit
- LFM2.5-Audio-1.5BLiquid AI1.5B≈1.0 GB at 4-bit
- Llama-OuteTTS-1.0-1BOuteAI1.2B≈0.8 GB at 4-bit
- stable-audio-open-1.0Stability AI1.2B≈0.8 GB at 4-bit
- VibeVoice-Realtime-0.5BMicrosoft1.0B≈0.7 GB at 4-bit
- csm-1bsesame1.0B≈0.7 GB at 4-bit1 also selling it hosted
- metavoice-1B-v0.1MetaVoice1.0B≈0.7 GB at 4-bit
- indic-parler-ttsAI4Bharat0.9B≈0.6 GB at 4-bit
- Qwen3-TTS-12Hz-0.6B-CustomVoiceAlibaba0.9B≈0.6 GB at 4-bit
- VoxCPM1.5OpenBMB0.8B≈0.5 GB at 4-bit
- neutts-airNeuphonic0.7B≈0.5 GB at 4-bit
- parler_tts_mini_v0.1Parler TTS0.6B≈0.4 GB at 4-bit
- Audio8-TTS-Preview-0.6bAudio80.6B≈0.4 GB at 4-bit1 also selling it hosted
- Qwen3-TTS-12Hz-0.6B-BaseAlibaba0.6B≈0.4 GB at 4-bit
- MiraTTSYatharthS0.5B≈0.3 GB at 4-bit
- Fun-CosyVoice3-0.5B-2512QwenAudio0.5B≈0.3 GB at 4-bit
- VoxCPM-0.5BOpenBMB0.5B≈0.3 GB at 4-bit
- IndicF5AI4Bharat0.4B≈0.2 GB at 4-bit
- VoiceCraftpyp10.3B≈0.2 GB at 4-bit
- Audio8-TTS-Preview-0.1bEdge00.2B≈0.1 GB at 4-bit
- Soprano-1.1-80Mekwek0.1B≈0.1 GB at 4-bit