Pass IndexThe State of AISign in

step-audio-r1.5

StepFun's reasoning-oriented end-to-end speech model, an upgrade of R1.1 built for deep sound understanding: it thinks while speaking, performs deep inference and keeps interaction natural, with reinforcement learning used to improve long-form dialogue.

text + audio → text + audio · made by StepFun

$1.48$15.58per Mtok in / out
StepFun

Sold by 1 way

SellerLaneRate
StepFun apistandard$1.48per Mtok in$15.58per Mtok out$0.3per Mtok cachedplatform.stepfun.com · read 2026-08-25

About

step-audio-r1.5 — a text and audio model from StepFun, sold by one company from $1.48 in and $15.58 out per million tokens.

It takes text and audio and returns text and audio. The catalogue files it under speak. Only StepFun sells it, at $1.48 in and $15.58 out per million tokens.

Every current figure

Maker
StepFun
Register
model
Takes
text + audio
Returns
text + audio
Sellers
1
Price
$1.48 in and $15.58 out per million tokens — StepFun