Qwen3-Omni-30B-A3B-Captioner
Since the research community currently lacks a general-purpose audio captioning model, we fine-tuned Qwen3-Omni-30B-A3B to obtain Qwen3-Omni-30B-A3B-Captioner, which produces detailed, low-hallucination captions for arbitrary audio inputs.
text + image → text + image · made by Alibaba
Sold by
Nobody in the catalogue publishes a price for this yet.
About
Qwen3-Omni-30B-A3B-Captioner — a text and images model from Alibaba.
It takes text and images and returns text and images. It was published in September 2025. The catalogue files it under image. Its weights are published under the other licence, which attaches conditions the plain open licences do not.
Every current figure
- Maker
- Alibaba
- Register
- model
- Takes
- text + image
- Returns
- text + image
- Published
- September 2025
- Parameters
- 32 billion · read from its own weights
- Licence
- other
Known as 1 name
Qwen/Qwen3-Omni-30B-A3B-Captioner