VibeVoice-ASR
VibeVoice-ASR is a unified speech-to-text model designed to handle 60-minute long-form audio in a single pass, generating structured transcriptions containing Who (Speaker), When (Timestamps), and What (Content), with support for Customized Hotwords and over 50 languages.
audio → text · made by Microsoft
Sold by
Nobody in the catalogue publishes a price for this yet.
About
VibeVoice-ASR — a text model from Microsoft.
It takes audio and returns text. It was published in January 2026. The catalogue files it under transcribe. Its weights are published under MIT, so you may run it on your own machine, or buy it from whoever serves it cheapest.
Every current figure
- Maker
- Microsoft
- Register
- model
- Takes
- audio
- Returns
- text
- Published
- January 2026
- Parameters
- 8.7 billion · read from its own weights
- Licence
- mit
Known as 1 name
microsoft/VibeVoice-ASR