step-audio-r1.5
StepFun's reasoning-oriented end-to-end speech model, an upgrade of R1.1 built for deep sound understanding: it thinks while speaking, performs deep inference and keeps interaction natural, with reinforcement learning used to improve long-form dialogue.
text + audio → text + audio · made by StepFun
$1.48→$15.58per Mtok in / out
StepFun
Sold by 1 way
| Seller | Lane | Rate |
|---|---|---|
| StepFun api | standard | $1.48per Mtok in$15.58per Mtok out$0.3per Mtok cachedplatform.stepfun.com · read 2026-08-25 |
About
step-audio-r1.5 — a text and audio model from StepFun, sold by one company from $1.48 in and $15.58 out per million tokens.
It takes text and audio and returns text and audio. The catalogue files it under speak. Only StepFun sells it, at $1.48 in and $15.58 out per million tokens.
Every current figure
- Maker
- StepFun
- Register
- model
- Takes
- text + audio
- Returns
- text + audio
- Sellers
- 1
- Price
- $1.48 in and $15.58 out per million tokens — StepFun