Conditionally licensed speech models
29 in the catalogue today. Every one with what it costs, who sells it and where it stands.
Models that read text aloud. They are usually priced per character, which makes the arithmetic simple: a thousand characters is roughly a paragraph. The real choices are the voice itself, whether you may clone one, how much control you have over emotion and pacing, and latency — a model that sounds wonderful in a rendered file may be too slow to hold a conversation.
Weights you can download, under a licence that asks for something in return: an acceptable-use policy, a naming requirement, a revenue ceiling above which you must ask, a restriction on training other models. The Llama and Gemma families are the familiar cases. Read the actual licence before you build on one of these — the condition is usually easy to meet and occasionally fatal, and which it is depends on your product rather than on the model.
- Audio8-TTS-Preview-0.1bEdge0
- Breeze-TTS-2BreezeBlue
- DramaboxResemble AI
- HunyuanVideo-FoleyTencent Hunyuan
- IndexTTS-2.5Index Team
- LFM2-Audio-1.5BLiquid AI
- LFM2.5-Audio-1.5BLiquid AI
- MARS5-TTSCAMB-AI
- MisoTTSMiso Labs
- Qwen3 Omni 30B A3B InstructAlibaba$1.8→$6.9per Mtok in / out2 selling
- SonicCartesia
- XTTS-v1Coqui.ai
- XTTS-v2Coqui.ai
- higgs-tts-2-3b-baseBoson AI
- higgs-tts-3-4bBoson AI
- kani-tts-2-ennineninesix
- kani-tts-370mnineninesix
- magpie_tts_multilingual_357mNVIDIA
- personaplex-7b-v1NVIDIA
- s2-proFish Audio$15per Mtok in1 selling
- scenema-audioScenemaAI
- seamless-expressivefacebook
- stable-audio-open-1.0Stability AI
- stable-audio-open-smallStability AI
- supertonicSupertone
- supertonic-2Supertone
- supertonic-3Supertone
- tada-1bHume AI
- tada-3b-mlHume AI