Pass IndexThe State of AISign in

Qwen3-Omni-30B-A3B-Captioner

Since the research community currently lacks a general-purpose audio captioning model, we fine-tuned Qwen3-Omni-30B-A3B to obtain Qwen3-Omni-30B-A3B-Captioner, which produces detailed, low-hallucination captions for arbitrary audio inputs.

text + image → text + image · made by Alibaba

Sold by

Nobody in the catalogue publishes a price for this yet.

About

Qwen3-Omni-30B-A3B-Captioner — a text and images model from Alibaba.

It takes text and images and returns text and images. It was published in September 2025. The catalogue files it under image. Its weights are published under the other licence, which attaches conditions the plain open licences do not.

Every current figure

Maker
Alibaba
Register
model
Takes
text + image
Returns
text + image
Published
September 2025
Parameters
32 billion · read from its own weights
Licence
other

Known as 1 name

Qwen/Qwen3-Omni-30B-A3B-Captioner