Pass IndexThe State of AISign in

Qwen3 Omni 30B A3B Instruct

The Qwen-Omni model accepts combined inputs of text and a single additional modality (image, audio, or video) to generate responses in text or speech.

text + image + audio + video → text + audio · made by Alibaba

$0.25$0.97per Mtok in / out
Novita AI · 2 sellers

Sold by 2 ways

SellerLaneRate
Novita AI aggregatorstandard$0.25per Mtok in$0.97per Mtok outapi.novita.ai · read 2026-08-25
PPIO aggregatorstandard$1.8per Mtok in$6.9per Mtok outapi.ppinfra.com · read 2026-09-16

About

Qwen3 Omni 30B A3B Instruct — a text and audio model from Alibaba, sold by 2 companies from $0.25 in and $0.97 out per million tokens.

It takes text, images, audio and video and returns text and audio, with a context window of 65,536 tokens. It was published in September 2025, trained on material up to April 2024. Its sellers say it can call a tool. The catalogue files it under speak. Two companies sell it. The cheapest is $0.25 in and $0.97 out per million tokens at Novita AI.

Every current figure

Maker
Alibaba
Register
model
Takes
text + image + audio + video
Returns
text + audio
Context
65,536 tokens
Longest answer
16,384 tokens
Published
September 2025
Knowledge to
April 2024
Parameters
30 billion · read from its own name, the total of a mixture
Licence
not read
Sellers
2
Maker's own price
not read
Price
$0.25 in and $0.97 out per million tokens — Novita AI

Known as 1 name

qwen/qwen3-omni-30b-a3b-instruct