Conditionally licensed transcription models
3 in the catalogue today, and the list grows as the market does. Every one with what it costs, who sells it and where it stands.
Models that turn recorded speech into text. They are metered by the minute or the second of audio, so cost follows the length of the recording and not the difficulty of it. What separates them is languages covered, whether they mark who is speaking, and whether they run in real time or only on a finished file — a model that is excellent on a podcast may be unusable on a live call.
Weights you can download, under a licence that asks for something in return: an acceptable-use policy, a naming requirement, a revenue ceiling above which you must ask, a restriction on training other models. The Llama and Gemma families are the familiar cases. Read the actual licence before you build on one of these — the condition is usually easy to meet and occasionally fatal, and which it is depends on your product rather than on the model.
- NVIDIA Nemotron 3.5 ASR Streaming 0.6BNVIDIA
- SenseVoiceSmallQwenAudio
- nemotron-speech-streaming-en-0.6bNVIDIA