smol-vision
Latest examples 👇🏻 - Fine-tune ColPali for Multimodal RAG - Fine-tune Gemma-3n for all modalities (audio-text-image) - Any-to-Any (Video) RAG with OmniEmbed and Qwen.
text + image → text · made by merve
Sold by
Nobody in the catalogue publishes a price for this yet.
About
smol-vision — a text model from merve.
It takes text and images and returns text. It was published in July 2025. The catalogue files it under chat.
Every current figure
- Maker
- merve
- Register
- model
- Takes
- text + image
- Returns
- text
- Published
- July 2025
Known as 1 name
merve/smol-vision