step-audio-2
StepFun's end-to-end speech model with broad auditory understanding across Mandarin, dialects, English and Japanese.
text + audio → text + audio · made by StepFun
$1.48→$10.39per Mtok in / out
StepFun
Sold by 1 way
| Seller | Lane | Rate |
|---|---|---|
| StepFun api | standard | $1.48per Mtok in$10.39per Mtok out$0.3per Mtok cachedplatform.stepfun.com · read 2026-08-25 |
About
step-audio-2 — a text and audio model from StepFun, sold by one company from $1.48 in and $10.39 out per million tokens.
It takes text and audio and returns text and audio. The catalogue files it under speak and search. Only StepFun sells it, at $1.48 in and $10.39 out per million tokens.
Every current figure
- Maker
- StepFun
- Register
- model
- Takes
- text + audio
- Returns
- text + audio
- Sellers
- 1
- Price
- $1.48 in and $10.39 out per million tokens — StepFun