Pass IndexThe State of AISign in

smol-vision

Latest examples 👇🏻 - Fine-tune ColPali for Multimodal RAG - Fine-tune Gemma-3n for all modalities (audio-text-image) - Any-to-Any (Video) RAG with OmniEmbed and Qwen.

text + image → text · made by merve

Sold by

Nobody in the catalogue publishes a price for this yet.

About

smol-vision — a text model from merve.

It takes text and images and returns text. It was published in July 2025. The catalogue files it under chat.

Every current figure

Maker
merve
Register
model
Takes
text + image
Returns
text
Published
July 2025

Known as 1 name

merve/smol-vision