Pass IndexThe State of AISign in

stepaudio-2.5-tts

StepAudio 2.5 TTS is StepFun's contextual speech-synthesis model that combines global and inline context control with zero-shot voice cloning, so emotion, pacing, pauses and delivery are described in plain natural language rather than tags.

text + audio → audio · made by StepFun

$0.000086per character
StepFun

Sold by 1 way

SellerLaneRate
StepFun apistandard$86per million charactersplatform.stepfun.com · read 2026-08-25

About

stepaudio-2.5-tts — an audio model from StepFun, sold by one company from $86 per million characters.

It takes text and audio and returns audio. It was published in April 2026. The catalogue files it under speak. Only StepFun sells it, at $86 per million characters.

Every current figure

Maker
StepFun
Register
model
Takes
text + audio
Returns
audio
Published
April 2026
Licence
not read
Sellers
1
Price
$86 per million characters — StepFun