stepaudio-2.5-tts
StepAudio 2.5 TTS is StepFun's contextual speech-synthesis model that combines global and inline context control with zero-shot voice cloning, so emotion, pacing, pauses and delivery are described in plain natural language rather than tags.
text + audio → audio · made by StepFun
$0.000086per character
StepFun
Sold by 1 way
| Seller | Lane | Rate |
|---|---|---|
| StepFun api | standard | $86per million charactersplatform.stepfun.com · read 2026-08-25 |
About
stepaudio-2.5-tts — an audio model from StepFun, sold by one company from $86 per million characters.
It takes text and audio and returns audio. It was published in April 2026. The catalogue files it under speak. Only StepFun sells it, at $86 per million characters.
Every current figure
- Maker
- StepFun
- Register
- model
- Takes
- text + audio
- Returns
- audio
- Published
- April 2026
- Licence
- not read
- Sellers
- 1
- Price
- $86 per million characters — StepFun